1Cleaning
Seven steps, in order. Each opens on an analysis screen showing how many rows are affected and where, before you fix anything. Nothing is written to your corpus until you apply.
Bibliometric tooling
A Scopus export is not a dataset. Identifiers are missing, venues are blank, affiliations stop at the institution name, and references are free text no co-citation analysis can read. Bibliominer walks you through fixing all of it, and records what it changed, so every number you publish can be justified.
Free, in your browser. Nothing is written until you apply.
Every term below is a function of the package, not a promise. No figures here on purpose: numbers belong to your corpus, not to a landing page.
Corpus · Actors · Impact · Concepts · Networks
Three pieces
Seven steps, in order. Each opens on an analysis screen showing how many rows are affected and where, before you fix anything. Nothing is written to your corpus until you apply.
Indicators, co-authorship and co-citation networks, figures. It reads the cleaned corpus and reaches the internal unit and the city, where a raw export stops at the institution name.
The whole computation lives in two Python packages. The apps only ask and draw, so anything they can do, your own scripts can do too.
The pipeline
The order is not decorative. The DOI is what finds the metadata, the metadata is what tells whether a venue is worth looking up, the authors are what build the affiliation grid, and references only reconcile against a corpus that is finally reliable.
Find the identifier of documents Scopus exported without one.
Everything downstream keys on the DOI: without it a document cannot be matched against Crossref or OpenAlex, and its references stay raw text.
Fill in year, citations, document type and language.
A missing year removes the document from every time series; a missing type lets an editorial count as a research paper.
Name the publication venues Scopus left blank.
A venue with no name cannot be ranked, compared or counted, and the documents it carries silently drop out of every source-level figure.
Restore authors and map each one to its affiliation.
Author positions are preserved across the three Scopus columns, so a collaboration network can be rebuilt without guessing who wrote with whom.
Resolve institution, city and country for every affiliation.
This is what lets the analysis reach the internal unit and the city, where a raw export stops at the institution name.
Restore titles, abstracts and keywords.
These are the only columns a thematic analysis reads. An empty abstract is a document that no topic model will ever see.
Reconcile every cited reference against Crossref and OpenAlex.
A reference is kept only when corroborated. A single unconfirmed source would fabricate co-citation links that do not exist.
What it guarantees
Every suggestion shows what it was matched on, and you accept or reject it. Automatic merging would confuse two institutions that merely share a name, and the error would be undetectable in the final corpus.
Removed documents are written to a file with the reason. A corpus you cannot account for is a corpus you cannot defend.
Download the corpus at any point and reimport it later to pick up where you left off. Your work never depends on a session staying open.