PaperAtlas

PaperAtlas is an atlas of the computational biology tools described in the open-access biomedical literature, built in a single automated pass over the PubMed Central full-text archive.

By the numbers

Articles screened6,446,741
Papers where computation is a central result1,074,191
Structured records extracted1,055,889
Method categories in the atlas910
Papers in the atlas155,587
Distinct tool names112,621
Of those, listed in none of bio.tools, Bioconductor, CRAN or PyPI85,955 (76.3%)
Dataset accessions recovered499,531
Code repositories recovered152,097

The atlas is free to browse. The published files are the July 2026 release, the version described in the manuscript. Releases are versioned by date and earlier releases are kept.

How it is built

The PaperAtlas pipeline: JATS XML parsing, abstract screening, guided-JSON full-text extraction, embedding and density-based clustering, restriction to biomedical topics, and static-site serving; with the acquisition funnel, publication-year distribution and screening outcome.
The pipeline and the corpus. Every article in the PubMed Central bulk archive is parsed, every abstract is screened for whether computation is a central result, and a structured record is extracted from each retained full text. The funnel runs from 7,283,594 articles parsed to the 155,587 retained in the atlas.
The atlas: a UMAP projection of tool-paper embeddings coloured by parent topic, the cluster-size distribution, EDAM ontology alignment per cluster, and registry coverage against bio.tools, PyPI, CRAN and Bioconductor.
The atlas and its coverage. Papers that contribute a named tool are embedded and clustered into 910 method categories, each of which matches a term in the EDAM ontology. Matching the 112,621 tool names against four registries leaves 85,955 of them (76.3%) in none.

Citation

Shah N, Vippagunta A, Bhargava Y. PaperAtlas: an automatically constructed atlas of computational biology tools from 6.4 million open-access articles. 2026.