Guidelines
How to use it well
Bibliominer is deliberately slower than a script that cleans everything at once. The reason is simple: a corpus you cannot account for is a corpus you cannot defend. Here is what the tool holds itself to, and what it expects from you.
The two guides
Every screen, every rule and every figure explained, with screenshots.
Cleaning guide
A real 128-article corpus, step by step: the problem as it is in the imported file, what Bibliominer does on its own (Crossref, OpenAlex, SCImago, GeoNames, ROR), what you decide, and the result.
Download the PDFAnalysis guide
Every screen and every chart of the analysis, with the exact formula behind each value and how to read it.
Download the PDFThe analysis package
The whole computation lives in the packages. The web applications only load a corpus, apply a filter and call them, an indicator written in an API would be invisible to everyone using the library, so nothing is. Anything the apps can do, your own scripts can do too.
bibliominer-analysis
Bibliometric analysis of a cleaned Scopus corpus: indicators, co-authorship and co-citation networks, figures.
pip install bibliominer-analysisversion 0.1.0 · Python >=3.9 · not published yet
Analysis, from a script
Corpus is the single entry point. It reads the cleaned CSV into six linked tables, and every indicator is computed from them.
from bibliominer_analysis import Corpus
corpus = Corpus.from_csv("corpus_cleaned.csv")
print(corpus.documents) # the corpus, as tablesThe package is laid out by concern: io/ reads the CSV into six tables linked by eid, metrics/ holds the indicators, one module per family , networks/ builds and measures the graphs, and model/ exposes Corpus.
Cleaning stays an interface
Cleaning is done in the web application, not from pip: every step asks for your decision, and the interface is where those decisions are made and recorded. Only the analysis is distributed as a package.
The five principles
Nothing is decided for you
Every suggestion shows what it was matched on, and you accept or reject it. Automatic merging would confuse two institutions that merely share a name, and the error would be undetectable in the final corpus.
Nothing disappears silently
Removed documents are written to a file with the reason. A corpus you cannot account for is a corpus you cannot defend.
The file is the save
Download the corpus at any point and reimport it later to pick up where you left off. Your work never depends on a session staying open.
Your keys, your quota
GeoNames, OpenAlex and Semantic Scholar are called with your own credentials. They stay private to you, and no step ever runs on someone else's account.
Every number is traceable
The final export ships with a cleaning report: where the corpus started, what was removed and why, what was recovered and with which tool. It is the Methods section of your paper.
What the tool refuses
Some steps will not let you continue. They are not obstacles for their own sake, each one blocks a corpus that would produce figures nobody can trust.
Incomplete metadata
Year, citations, document type and language must all be filled before Sources. A missing year removes the document from every time series without saying so.
Authors without affiliation
You either complete the article or remove it explicitly, and the removal is logged. An article with no author distorts every collaboration and productivity figure.
Merging on an unfinished city
Organisation names are merged on how often each spelling is used. Those counts are not final until every affiliation of the city is resolved, so the merge waits.
Exporting before reconciliation
The final file replaces the References column with the reconciled list. Built too early, it would carry raw Scopus text, which no co-citation analysis can read.
Before you publish your numbers
Four things worth checking, in order of how often they go wrong.
Read the removals
removed_articles.csv lists every document that left the corpus, with its reason. If a count surprises you later, this is the first file to open.
Check the merges you confirmed
Each merge rewrote an organisation name across the whole corpus. The list of confirmed merges stays available, and each one can be undone, the original names come back.
Look at what was resolved for you
Fixing one article often makes others consistent as a side effect. Those are marked resolved by another fix, nobody reviewed them. Open a few.
Keep the cleaning report
It states where the corpus started, what was removed and why, what was recovered and with which tool. It is what makes the manipulation reproducible.
API keys
Three services are called during cleaning, each with your own credentials: GeoNames for cities and countries, OpenAlex for DOI recovery and reference reconciliation, and Semantic Scholar, optional, roughly doubling the run for a marginal gain.
They are free to obtain and stay private to you. Without them the matching steps stay locked rather than running on someone else’s quota. You can register several keys per service: when one hits its daily limit, the run rotates to the next instead of stopping.
Working on the code
One Python environment serves the whole project, at the repository root. It sits there rather than inside one of the pieces precisely because it belongs to none of them: the API and the packages depend on it equally.
.\setup.ps1 # build the environment, install everything
.\setup.ps1 -Force # rebuild from scratch
.\run.ps1 test # run the package testsThe packages are installed in editable mode: they point at their sources, so a change is picked up without reinstalling. A frozen copy in site-packages would import just as well, and you would be testing a version you are no longer editing.