Preview: 3 identity value(s) are still placeholders. Names, affiliation and contact are filled in src/site.config.ts before publishing.

How it works

Seven steps, then the file

Each step opens on an analysis screen: how many rows are affected, of what kind, and where. You read before you fix, and nothing is written to the corpus until you apply.

The pipeline at a glance
  1. ScopusYour Scopus exportCSV, as downloaded
  2. Bibliominer Cleaningin your browser, step by step
  3. A cleaned corpusCSV + before/after report
  4. Bibliominer Analysisindicators, networks, maps

Where each answer comes from

  1. 1DOI
  2. 2MetadataYou
  3. 3SourcesSCImago
  4. 4Authors
  5. 5AffiliationsGeoNames
  6. 6TextCorpus vocabulary
  7. 7References

You decide: every suggestion is approved or refused by you.

The order, and why it is the order

The DOI is what finds the metadata. The metadata is what tells whether a venue is worth looking up. The authors are what build the affiliation grid. And the references only reconcile against a corpus that is finally reliable. Reconciling first would spend API quota on documents that are still wrong.

1

DOI

Find the identifier of documents Scopus exported without one.

Everything downstream keys on the DOI: without it a document cannot be matched against Crossref or OpenAlex, and its references stay raw text.

2

Metadata

Fill in year, citations, document type and language.

A missing year removes the document from every time series; a missing type lets an editorial count as a research paper.

3

Sources

Name the publication venues Scopus left blank.

A venue with no name cannot be ranked, compared or counted, and the documents it carries silently drop out of every source-level figure.

4

Authors

Restore authors and map each one to its affiliation.

Author positions are preserved across the three Scopus columns, so a collaboration network can be rebuilt without guessing who wrote with whom.

5

Affiliations

Resolve institution, city and country for every affiliation.

This is what lets the analysis reach the internal unit and the city, where a raw export stops at the institution name.

6

Text

Restore titles, abstracts and keywords.

These are the only columns a thematic analysis reads. An empty abstract is a document that no topic model will ever see.

7

References

Reconcile every cited reference against Crossref and OpenAlex.

A reference is kept only when corroborated. A single unconfirmed source would fabricate co-citation links that do not exist.

What you get at the end

The cleaned corpus

A CSV in Scopus format, with two Bibliominer conventions: author positions preserved across the three author columns, and each affiliation segment carrying its own label.

A before / after report

Empty cells per column, before and after. Articles removed and why. Affiliations consolidated. References Scopus → Global. These are the numbers you quote in a Methods section.

The reconciliation tables

One row per article, one row per cited reference with its provenance. They exist nowhere else, and the API calls behind them have already been spent.

The contract between cleaning and analysis

The cleaned file is not just tidier. It carries structure a raw export does not have. That structure is what lets the analysis reach the internal unit and the city.

Author positions are preserved, so the three columns stay aligned and a collaboration network can be rebuilt without guessing who wrote with whom:

Authors            1:Idri A.; 2:Hosni M.; 3:Abran A.
Author full names  1:Idri, Ali (6602789810); 2:Hosni, Mohamed (57189341317)
Author(s) ID       1:6602789810; 2:57189341317; 3:7004233119

Affiliation segments are labelled, so an empty field is simply omitted and reading never depends on position:

subparent: ENSMR, parent 1: National School of Mineral Industry, city: Rabat, country: Morocco
parent 1: University of Murcia, city: Murcia, region: Murcia, country: Spain
parent 1: LIRIMA, site: virtual, country: France

subparent, parent 1, parent 2, city, site, region, country. Several affiliations of one article stay separated by ; .

Bibliometric analysis

Once the corpus holds, measure it

Cleaning is the means; this is the point. Five families of indicators, each answering one question, and each stating what it does NOT let you say.

Five sections, five questions

  1. 1Corpus

    What does this corpus contain, and what can it support?

    Documents, period, sources, annual growth rate, trajectory over time. The figures that tell whether the corpus is large enough for what you intend to ask of it.

  2. 2Actors

    Who produces, authors, institutions, journals, cities, countries

    Rankings, output over time, documents by number of signatories. This is the level that only exists because affiliations were resolved down to the city.

  3. 3Impact

    What counts, and for whom

    Global citations against local ones, the historiograph, RPYS spectroscopy. A paper with 800 global citations and none local matters elsewhere, not here.

  4. 4Concepts

    What this corpus is about, and how it shifts

    Word clouds and treemaps, topics and their position in time, term dynamics. Two raw materials that must never be confused: author keywords and indexed keywords.

  5. 5Networks

    Who works with whom, which ideas travel together

    Six families of networks, co-authorship, co-citation, bibliographic coupling, co-word. Two safeguards, without which none of them can be read.

Three things to know before reading any number

1

Global citations are not local citations

The most useful distinction in bibliometrics, and the most misread. Global counts what Scopus counts, all sources together. Local counts how many documents OF THIS CORPUS cite it. A paper with 800 global and 0 local is important elsewhere, not here, the reverse signals foundational work for the community you are studying.

2

A filter applies to everything

Years, document types, countries, the filter bar applies to every figure on every page. A number read on a filtered page does not describe the whole corpus, and nothing on the figure itself will remind you of that.

3

Coverage decides validity

The Quality page states, analysis by analysis, whether the field it depends on is filled enough. An h-index computed on a corpus missing 40% of its citations is a number, not a measurement. Start there.