Preview: 3 identity value(s) are still placeholders. Names, affiliation and contact are filled in src/site.config.ts before publishing.

Bibliometric tooling

From a raw Scopus export
to a corpus you can defend.

A Scopus export is not a dataset. Identifiers are missing, venues are blank, affiliations stop at the institution name, and references are free text no co-citation analysis can read. Bibliominer walks you through fixing all of it, and records what it changed, so every number you publish can be justified.

Free, in your browser. Nothing is written until you apply.

What the analysis contains

Every term below is a function of the package, not a promise. No figures here on purpose: numbers belong to your corpus, not to a landing page.

Lotkaco-citationh-indexBradfordlocal citationsco-authorshipZipfhistoriographg-indexthematic mapbibliographic couplingco-wordRPYSCAGRtrend topicsGinii10Lorenzm-indexco-institutionPrice indexdoubling timeco-countrycollaboration index

Corpus · Actors · Impact · Concepts · Networks

Three pieces

A cleaning app, an analysis app, and the libraries behind them

1Cleaning

Seven steps, in order. Each opens on an analysis screen showing how many rows are affected and where, before you fix anything. Nothing is written to your corpus until you apply.

2Analysis

Indicators, co-authorship and co-citation networks, figures. It reads the cleaned corpus and reaches the internal unit and the city, where a raw export stops at the institution name.

3Libraries

The whole computation lives in two Python packages. The apps only ask and draw, so anything they can do, your own scripts can do too.

The pipeline

Seven steps, and why each one exists

The order is not decorative. The DOI is what finds the metadata, the metadata is what tells whether a venue is worth looking up, the authors are what build the affiliation grid, and references only reconcile against a corpus that is finally reliable.

1

DOI

Find the identifier of documents Scopus exported without one.

Everything downstream keys on the DOI: without it a document cannot be matched against Crossref or OpenAlex, and its references stay raw text.

2

Metadata

Fill in year, citations, document type and language.

A missing year removes the document from every time series; a missing type lets an editorial count as a research paper.

3

Sources

Name the publication venues Scopus left blank.

A venue with no name cannot be ranked, compared or counted, and the documents it carries silently drop out of every source-level figure.

4

Authors

Restore authors and map each one to its affiliation.

Author positions are preserved across the three Scopus columns, so a collaboration network can be rebuilt without guessing who wrote with whom.

5

Affiliations

Resolve institution, city and country for every affiliation.

This is what lets the analysis reach the internal unit and the city, where a raw export stops at the institution name.

6

Text

Restore titles, abstracts and keywords.

These are the only columns a thematic analysis reads. An empty abstract is a document that no topic model will ever see.

7

References

Reconcile every cited reference against Crossref and OpenAlex.

A reference is kept only when corroborated. A single unconfirmed source would fabricate co-citation links that do not exist.

What it guarantees

You can account for every row

Nothing is decided for you

Every suggestion shows what it was matched on, and you accept or reject it. Automatic merging would confuse two institutions that merely share a name, and the error would be undetectable in the final corpus.

Nothing disappears silently

Removed documents are written to a file with the reason. A corpus you cannot account for is a corpus you cannot defend.

The file is the save

Download the corpus at any point and reimport it later to pick up where you left off. Your work never depends on a session staying open.