ArticleAtlas

A map of every arXiv preprint, coloured by which subjects are heating up.

Every arXiv preprint placed by what it is about. Papers that read alike sit together, each cluster is named, and the colour shows which subjects are growing. A cluster is a set of papers on the same subject, not an arXiv category. The map is free to browse and the data behind it is free to download.

How to read it. A blob's size is how many papers the cluster holds; its colour is how fast that cluster is growing against arXiv as a whole. So a large grey blob is a big settled subject, a small red one is a subject taking off, and a large blue one is a subject going quiet even though it still publishes plenty.

Full screen ↗

Zoom in for detail: clusters first, then individual papers. Pick a cluster — on the map or in the panel — to see its recent papers and how its publishing rate has changed. Search jumps straight to a cluster. Topics groups papers by what they are about; Categories shows arXiv's own categories instead.

Version 2026.09.20, built 2026-09-20, from the arXiv metadata harvest of 20 Sep 2026 02:08:33 UTC. CC-BY-4.0.

Papers
3,166,401
Clusters
425
Clustered
1,274,776
Unclustered
59.7%

The archive holds everything: map.json, the per-cluster paper lists (116.2 MB in 425 files), the monthly counts, the category graph and every paper's cluster assignment.

Prefer a single file? map.json (226 KB) · timeseries.json (1.4 MB) · assignments.csv.gz (5.9 MB) · or one cluster at a time from /data/topic-map/clusters/.

Look inside a cluster

Every cluster, and every arXiv category, with its publishing rate month by month and its newest papers. The same drill-in the map gives, without opening it.

Open in the map
Loading…

The files

FileOne row is
map.json A cluster: its name, a one-line description, how many papers it holds, where it sits on the map, its distinctive words, its main arXiv categories, and how fast it is growing. Also which clusters are close to each other, and where the data came from — version, build date and the arXiv harvest behind it.
clusters/<id>.json A paper of that cluster, newest first: its arXiv identifier, title, month of submission, categories and position on the map. One file per cluster. Where a cluster has been read for claims, its papers also carry what they say that no earlier paper in the same cluster had said.
timeseries.json Papers in one cluster in one month, back to 1991.
assignments.csv.gz One line per paper: which cluster it belongs to, for matching the map against your own data.

Every file names the dataset version it belongs to. A paper that is not close enough to any group of similar papers is left out of the cluster files and counted as unclustered above — that is a large share, and the map says so rather than hiding it.

Which subjects are heating up

A cluster's colour is its share of the last twelve months divided by its share of all time. Growing faster than arXiv as a whole is red, slower is blue. That is a measure of attention, not of quality or importance, and a cluster can grow simply because the subject is fashionable. Clusters too small or too young to judge are left grey. The current lists are here, and picking any cluster below shows its own curve.

Which papers say something new

This one is about a single paper, and it is not a rating. A language model reads the abstract and writes down what the paper claims — what it proposes, what it compares, what it measures. Each claim is then looked up among the claims of papers published earlier in the same cluster. What a paper gets is the number of claims nobody there had made yet, and the sentences they came from, so you can check the answer rather than trust it.

Reading abstracts is slow, so this has only been done for a few clusters so far, and within them only for recent papers. That has consequences worth knowing. The comparison is against the papers read so far, not against everything ever published, so a claim called new may already sit in a paper nothing has read — the number can only come down as more is read, never up. A paper with nothing shown has not been read; it does not mean nothing new was found. And a paper filed in the wrong cluster is compared against the wrong neighbours.

What this one does differently

Maps of the literature are not new, and the technique behind this one is standard. Four things are less usual:

What it is not: there is no citation data here, so it cannot tell you what cites what — tools built on citation graphs do that better. Clusters and their names come from the wording of titles and abstracts alone. More than half of all papers sit in no cluster and are not drawn; the exact share is in the panel above.

How it is built

Each title and abstract is turned into a numerical fingerprint of its meaning, papers with similar fingerprints are grouped into clusters, and the clusters are placed on the map so that related ones sit near each other. A language model then writes each cluster's name from the papers closest to its centre.

Source data

Only arXiv's own metadata: identifiers, titles, abstracts, categories and dates, from arXiv's public interface for bulk metadata. No citation data, no author or affiliation data, and no full text — paper licences vary, so this site never re-hosts the papers themselves.

Using the map elsewhere

Every view has its own link — the address bar updates as you move around — so the simplest way to point someone at what you are looking at is to copy it. The map can also be placed in another page; ask if you would like to.

Licence and citation

The dataset is CC-BY-4.0. The arXiv metadata it is built from is CC0.

To cite it, name the version — it identifies the metadata harvest the map was built from:

Morozov, D. ArticleAtlas: an atlas of the literature, coloured by what is heating up. articleatlas.org, dataset version (see above).

Get in touch

A cluster named badly, a number that looks wrong, a use of the dataset I should know about — write to [email protected]. There is no account system and nothing to sign up for; mail is the whole feedback channel.