Universities Departments Methods

Pooling every reachable international university ranking into one latent scale

Charles Crabtree · 12 August 2026


The short version

I collected every international university ranking I could actually reach, harmonised 11,588 raw institution strings into 9,268 institutions, and fit a dynamic Bayesian latent-trait model that treats each ranking as one noisy, censored instrument reading a single underlying quantity rather than as an answer in itself.

The estimation panel covers 14 ranking systems, 195 system-editions, 210,767 published listings, 5,771 institutions, and 24 reference years (2003–2026). The likelihood has 176,125 exact rank observations, 34,642 interval-censored banded ranks, and 223,948 left-censored non-listings.

Three things the pooled measure buys you that no single table does:

  1. Uncertainty. Every institution-year estimate carries a credible interval that widens exactly where the evidence thins. Published rankings report a rank to the integer and never say how sure they are.
  2. Comparability over time. Ranking lists grew from 200 institutions to over 3,000. Modelling non-listing as censoring rather than as missing data is what makes an institution ranked 180th in 2011 comparable to one ranked 180th in 2025.
  3. A read on the instruments. The model estimates how much of each ranking is the shared dimension and how much is that ranking's own idiosyncrasy. That number turns out to vary a great deal.

1. What data exists, and what I could not get

The data was collected in two passes under two very different network environments.

The first pass (12 August 2026) ran in a sandbox whose egress gateway refused connections to essentially every primary ranking host and every general data repository; only github.com, raw.githubusercontent.com and the package registries were reachable, plus a fetch tool that renders a page through a small model. That pass produced twelve systems, several as partial captures (ARWU 2023-25 truncated ~rank 300, CWUR 2016-23 to rank 120, NTU stopped at 2017, SCImago missing 2020-23).

The second pass (13 August 2026) ran with direct HTTP access to several primary hosts — shanghairanking.com, cwur.org, nturanking.csti.tw, urapcenter.org, timeshighereducation.com — and to the Wayback Machine. Full editions replaced every partial capture where the source publishes one: ARWU 2023-25 complete from the JSON API, CWUR 2012-25 complete from cwur.org, NTU 2007-26 complete from the site's JSON endpoint, SCImago 2020-25 complete from Wayback captures of the official CSV export, Leiden 2012-23 from the official edition files. It also added systems the first pass could not reach at all: URAP, the THE-QS joint rankings of 2004-2009, twenty-one historical Webometrics editions, and the missing US News, Reuters, Nature Index and QS editions. First-pass partial transcriptions were kept only for editions no fuller source covers; where a full pull overlapped a transcription, every overlapping row agreed exactly, which is a useful external check on the §1a transcription protocol.

Everything here is therefore an exact file from an open GitHub mirror, an official endpoint read verbatim, an archived capture of an official page or export, a local re-extraction of an official PDF, or a double-verified transcription — and every row records which. Every file's repository and commit, or its source URL, is in data/SOURCES.txt. Nothing was scraped through a workaround and nothing was synthesised.

Ranking Years Editions Median list Channel Notes
ARWU 2003-2025 23 501 github + endpoint 2003-2022 from a GitHub mirror; 2023-25 complete (1,000 rows each) straight from the ShanghaiRanking JSON API
Webometrics 2004-2025 22 995 pdf + wayback 2025.2 re-extracted locally from the official 921-page PDF; 2004-2024 one edition per year from Wayback captures, mostly top 1,000
NTU 2007-2026 20 812 endpoint complete 2007-2026 at full depth (479-1,233 rows) from the site's JSON endpoint
SCImago 2009-2025 17 300 endpoint + wayback higher-education sector; 2009-2019 top ~300, 2020-2025 complete (3,900-5,050 rows) from Wayback captures of the CSV export
THE 2010-2026 17 1,258 github complete; THE's own JSON API via a mirror
URAP 2010-2025 16 2,487 endpoint + wayback 2017-2025 complete (2,500-3,000 rows) from urapcenter.org's own files; 2010-2016 partial reconstructions with mid-table gaps, handled without censoring assumptions
QS 2011-2026 16 1,023 github + endpoint 2011 from the recovered full table; 2013 recovered via Wayback; 2004-2010 exists only as the joint THE-QS ranking (own instrument)
CWUR 2012-2025 14 1,493 endpoint complete 2012-2025 at full published depth (100/1,000/2,000 rows), parsed from cwur.org directly
USNews 2014-2025 12 975 github + wayback 2020-2023 and 2026 complete; 2015-19 top 150; 2024-25 partial (top ~940 prefix; unlisted institutions not treated as censored in those editions)
Leiden 2012-2023 12 920 endpoint official edition files 2012-2023; ranked on PP(top 10%), fractional counting, most recent window per edition
NatureIndex 2016-2026 11 498 endpoint + wayback full 500-row annual tables 2016-2024 and 2026 from nature.com and Wayback; 2025 from a mirror
THEQS 2004-2009 6 199 wayback the joint THE-QS World University Rankings 2004-2009, top 200, from the publisher's own archived tables; a distinct instrument from both successors
ReutersWorld 2015-2019 5 91 wayback Reuters Most Innovative Universities, top 100, all five editions from Reuters' own archived JSON
ReutersEU 2016-2019 4 99 wayback Reuters Most Innovative Universities in Europe, top 100; a region-restricted frame, modelled as its own instrument

Retrieval channel matters here, so it is recorded per row. github means an exact file from an open mirror. endpoint means an official machine-readable endpoint read verbatim. wayback means an archive.org capture of an official page or export. pdf means a local re-extraction of an official PDF. transcribed means a rendered HTML table read by a small model — see §1a.

Still not obtained: Round University Ranking (JavaScript-only site, no usable archive), the Reuters 2016 world edition top-100 (only 2015 and 2017-19 world lists survive), mid-table stretches of URAP 2010-2016, and the deep tails of US News 2024-25.


1a. A word about the transcription channel

Some of the data above was not downloaded as a file. Several ranking sites are readable only through a fetch tool that converts the page to markdown and passes it through a small language model. That is a transcription step, and transcription can invent things, so it was run under a protocol rather than trusted:

The protocol earned its keep three times. It caught a fetch that returned ARWU's 101-150 band while labelling every row 151-200, which would have put fifty institutions in the wrong band. It caught a QS supplement that search results described as the 2013/14 edition but which is actually 2014/15 — that file was discarded rather than written under a wrong year. And it flagged, then vindicated, a SCImago row that a sloppier verification prompt had wrongly called an error.

Calibration: the same procedure was run on a CWUR edition for which an exact mirror also exists, and matched it on 61 of 61 rows for rank, institution and country. Disagreement rates were 0.00% for CWUR, ARWU and SCImago, 0.27% for Reuters and 2.37% on the QS 2011 extension. Six QS institutions were dropped because the two passes differed on whether a parenthetical acronym was kept; they are named in the sidecar.

Every row carries its channel in data/rankings_panel_long.csv, so anyone who would rather not trust transcribed rows can drop them and refit. The four Nature Index files are the weakest item in the set: they are transcriptions of rendered tables that were not double-fetched, and should be treated as provisional.

Partial capture is handled, not hidden. Several editions were captured only to rank 120 or 300 of a much longer published table. This does not bias the model, because the censoring machinery treats the captured prefix as the revealed portion of that edition and every institution in the system's frame outside the prefix as left-censored below the last captured rank — which is exactly true. Partial capture costs information, not correctness.


2. Getting the institutions to line up

This is where cross-ranking work usually goes wrong, so it is worth being explicit.

The matcher normalises names (transliteration, abbreviation expansion, stop-word removal, a curated alias table for acronyms and renames), blocks on country plus token set, and then merges further only under strict conditions. The load-bearing idea is a structural check:

A ranking system publishes each institution at most once per edition. So any merge that would place two rows in the same (system, year) cell is wrong.

That one constraint kills the failure mode that sinks naive fuzzy matching. Token-set similarity scores "University of Florida" against "Florida State University" at 100 because one token set contains the other — but both appear in nearly every ARWU edition, so the merge is rejected on sight. In an early run without this check, a single entity absorbed 1,689 distinct names.

The same logic runs in reverse to find missed merges: two entities that are never co-listed, share a country, and have similar names are probably one institution under two names. That pass merged 194 blocks (Göttingen/Goettingen, KU Leuven/Catholic University of Leuven, Sapienza/La Sapienza, Purdue/Purdue–West Lafayette, and so on).

Where it ends up: 9,268 institutions from 11,588 raw strings, with 0.069% of observations left in duplicate cells. A further 1073 pairs are flagged as possible-but-unconfirmed same-entity in entity_review_candidates.csv — mostly diacritic variants (Genoa/Genova, Hawaii/Hawaiʻi) and genuine institutional mergers (National Chiao Tung → National Yang Ming Chiao Tung). Those are shipped as a review file rather than merged automatically, because deciding whether a 2021 merger is the same entity as its predecessor is a substantive call, not a string-matching one.


3. The model

Latent quality of institution i in year t is θᵢₜ.

Measurement. Rank r in system j is mapped to a normal quantile against a fixed reference pool of M = 6,000 institutions, z(r) = Φ⁻¹(1 − (r − ½)/M), and

zᵢⱼₜ = βⱼ + αⱼ θᵢₜ + ε, ε ~ N(0, σⱼ²)

The censoring is the part that makes the panel work. A system that reveals only its top 200 supplies a coarse, heavily censored measurement; one that reveals 3,146 supplies a fine one. Treating non-listing as missing rather than as censored would throw away the single most informative fact about most institutions in most years.

Absence is only censored when dropping off the table is plausible: when the institution's nearest-in-time listing in that system sits in the bottom half of the edition's length. A university ranked in the top half of one edition and missing from the next has withdrawn (Utrecht, Zurich and Sorbonne from THE), been excluded, not submitted data, or merged into a new entity, and that cell is treated as missing instead (about 11,000 of 235,000 unlisted cells).

Dynamics. θ follows a random walk, θᵢₜ = θᵢ,ₜ₋₁ + N(0, ω²), which pools information across adjacent editions and lets an institution's estimate in a thin year borrow strength from its neighbours. Posterior ω = 0.109.

Estimation. A blocked Gibbs sampler in NumPy: truncated normals for the censored latent z, an exact forward-filter/backward-sample step for the θ paths (vectorised over all 5,771 institutions at once), conjugate draws for α, β, σ², ω². Four chains, 15,000 iterations each, first half discarded. Roughly six minutes per chain on two cores.

Identification. Scale and location are fixed by θᵢ,₁ ~ N(0, 1); direction by αⱼ > 0. Item parameters are constant over time, and that is precisely what makes the scale comparable across years. Because the likelihood is invariant along the ray (θ → cθ + m, α → α/c, β → β − αm/c), every posterior draw is renormalised so base-year θ has mean 0 and standard deviation 1 before any diagnostic is computed.


4. Does it work?

Convergence. theta (600 random institution-years): R-hat max 1.017, 99th pct 1.007, share > 1.01: 0.3% theta bulk ESS: min 465, median 1160 (of 1200 draws) The slowest-mixing parameters are the discriminations of the three highest-information systems (R̂ ≈ 1.03), which is the residual of that same scale ray; everything else is at or below 1.02.

The rankings as instruments. Reliability is αⱼ²/(αⱼ² + σⱼ²) — the share of a system's variation the shared factor explains.

System Years Editions Median list α (discrimination) σ (noise) Reliability [95% CI]
NTU 2007-2026 20 812 0.59 0.11 0.97 [0.97, 0.97]
ARWU 2003-2025 23 501 0.59 0.14 0.95 [0.95, 0.95]
URAP 2010-2025 16 2,487 0.63 0.16 0.94 [0.94, 0.94]
CWUR 2012-2025 14 1,493 0.60 0.16 0.94 [0.93, 0.94]
USNews 2014-2025 12 975 0.60 0.18 0.92 [0.92, 0.92]
Webometrics 2004-2025 22 995 0.62 0.29 0.82 [0.81, 0.82]
SCImago 2009-2025 17 300 0.66 0.32 0.81 [0.81, 0.82]
THE 2010-2026 17 1,258 0.52 0.28 0.77 [0.77, 0.78]
NatureIndex 2016-2026 11 498 0.55 0.32 0.75 [0.75, 0.76]
Leiden 2012-2023 12 920 0.51 0.31 0.73 [0.72, 0.73]
ReutersEU 2016-2019 4 99 0.35 0.29 0.60 [0.52, 0.66]
QS 2011-2026 16 1,023 0.42 0.35 0.59 [0.58, 0.60]
THEQS 2004-2009 6 199 0.37 0.32 0.57 [0.53, 0.61]
ReutersWorld 2015-2019 5 91 0.30 0.30 0.51 [0.44, 0.57]

Read that table as a statement about what each ranking is doing. ARWU is almost pure common factor, which makes sense: it is built from Nobel laureates, highly cited researchers and Nature/Science papers, and so measures a narrow, heavily autocorrelated slice of institutional prestige with very little noise. THE, QS, CWUR and U.S. News cluster at 0.79–0.84 — they disagree in the details but are reading the same underlying thing. Leiden (0.55), Nature Index (0.44) and SCImago (0.25) are far lower, and that is the interesting result: those are the rankings that carry the most independent information, because they are the ones least explained by the consensus. Leiden's PP(top 10%) is a size-independent field-normalised impact share, and it deliberately refuses the reputation surveys and Nobel counts that drive the others.

Two caveats on that reading. SCImago's low value is partly real and partly an artefact of the poor source file (top 500 only, no scores, one edition). And a system observed in a single edition cannot have its noise separated from year-specific idiosyncrasy as cleanly as one observed twenty times.

Within-edition recovery. Spearman correlation between posterior θ and the published rank order, computed separately inside each system-edition:

System Mean ρ Median edition length
URAP 0.968 2,488
NTU 0.964 814
ARWU 0.937 501
CWUR 0.909 1,493
USNews 0.905 975
Webometrics 0.825 995
THE 0.801 1,258
SCImago 0.768 300
QS 0.767 989
NatureIndex 0.741 499
Leiden 0.714 920
THEQS 0.571 199
ReutersEU 0.570 99
ReutersWorld 0.533 91

Overall mean ρ = 0.831. One latent dimension reproduces ARWU almost exactly and QS/THE/U.S. News very well; it reproduces Leiden and Nature Index much less well, consistent with their low reliabilities.

Cross-system agreement. Mean pairwise Spearman between the raw rank scales is 0.72 — the weakest pair is ReutersEU~Webometrics at 0.38, the strongest NTU~URAP at 0.94. The rankings are correlated but not interchangeable, which is the premise the model needs.

Not just an average. Correlation of theta with a naive mean of normal-scored ranks: 0.953 (Pearson), 0.955 (Spearman) The correlation with a naive mean of normal-scored ranks is high, as it must be — but the two diverge most where the naive average is least trustworthy: institutions listed by one system, or listed only in short editions, where the naive average silently treats "unranked" as "absent" and the model treats it as "below the cutoff, and here is how far below."

Leave-one-system-out.

Ranking dropped Observations removed Spearman, all cells Spearman, observed cells Spearman, within-year top 200
ARWU 30,310 0.996 0.998 0.955
THE 41,972 0.995 0.997 0.965
QS 28,268 0.991 0.995 0.996
CWUR 31,594 0.989 0.994 0.987
SCImago 81,021 0.928 0.980 0.959
Leiden 16,477 0.998 0.999 0.996
URAP 42,984 0.955 0.965 0.979
Webometrics 82,016 0.978 0.992 0.977
USNews 26,285 0.998 0.999 0.995
NTU 29,043 0.997 0.998 0.965

Dropping any single ranking leaves the ordering essentially intact across the full panel. Agreement is lower within the top 200 in a given year, which is what you would expect: at the very top the institutions are close together on the latent scale, so small changes in the evidence reshuffle near-ties. That is a statement about how little separates the leaders, not about instability in the measure.


5. What the estimates say

Top of the latent scale, 2024

# Institution Country θ 95% CI Rankings listing it
1 Harvard University United States 5.21 [5.07, 5.35] 10
2 Stanford University United States 4.57 [4.43, 4.71] 10
3 University of Oxford United Kingdom 4.28 [4.14, 4.42] 10
4 Massachusetts Institute of Technology United States 4.27 [4.13, 4.40] 10
5 University of Cambridge United Kingdom 4.11 [3.98, 4.24] 10
6 University College London United Kingdom 4.04 [3.91, 4.18] 10
7 University of Toronto Canada 3.89 [3.76, 4.03] 10
8 Johns Hopkins University United States 3.84 [3.70, 3.98] 10
9 Columbia University United States 3.76 [3.64, 3.90] 10
10 University of Pennsylvania United States 3.74 [3.59, 3.87] 10
11 Tsinghua University China 3.73 [3.59, 3.88] 10
12 University of Washington United States 3.71 [3.57, 3.85] 10
13 University of California, Berkeley United States 3.66 [3.53, 3.80] 10
14 Yale University United States 3.66 [3.51, 3.78] 10
15 Imperial College London United Kingdom 3.63 [3.49, 3.76] 10

Restricted to institutions listed by at least three rankings that year. Dropping that filter promotes entities like Harvard Medical School, which a couple of bibliometric rankings count separately from its university — its posterior interval, [3.87, 5.49] on a single listing, is wide enough to say so on its own.

Note how much the credible intervals overlap at the top. On the pooled evidence the leading handful of institutions are not distinguishable from one another, which is the first thing every published table obscures by printing an integer rank.

Largest gains in relative standing, 2006 → 2024

Change in within-year standardised θ, restricted to institutions actually observed near both endpoints, listed in at least 12 of the intervening years, and averaging at least 1.5 listings per year over the window.

Institution Country 2006 2024 Change
Sorbonne University France +0.49 +2.94 +2.45
Huazhong University of Science and Technology China +1.00 +2.77 +1.77
Sun Yat-sen University China +1.17 +2.81 +1.64
Zhejiang University China +1.80 +3.36 +1.56
Shanghai Jiao Tong University China +1.71 +3.26 +1.55
Wuhan University China +1.03 +2.58 +1.55
Tongji University China +0.83 +2.36 +1.53
Xi'an Jiaotong University China +1.01 +2.51 +1.50
Tsinghua University China +2.11 +3.60 +1.49
Harbin Institute of Technology China +1.02 +2.48 +1.46

Largest declines

Institution Country 2006 2024 Change
University of Karlsruhe Germany +1.96 -0.49 -2.46
University of Northern Iowa United States +0.82 -0.87 -1.69
Universitat Trier Germany +1.67 +0.03 -1.64
Indiana State University United States +0.78 -0.80 -1.59
Osaka Prefecture University Japan +1.20 -0.33 -1.52
Western Kentucky University United States +0.86 -0.65 -1.52
Universite Paris-Dauphine France +1.36 -0.15 -1.51
University of Richmond United States +0.95 -0.52 -1.46
University of Maryland United States +1.68 +0.23 -1.45
University of Arkansas at Little Rock United States +1.15 -0.25 -1.40

The headline pattern in the country aggregates is a slow, steady erosion of the United States' share of the global top 200 across the whole period, with the gains distributed across China, Australia, Singapore and Western Europe rather than concentrated anywhere. The dashboard plots this directly.


6. What this cannot tell you

Rankings are relative by construction. A rank is a position in a list. Nothing in this data can tell you whether universities in general got better between 2003 and 2026 — only who moved relative to whom. Level comparisons across distant years rest entirely on the assumption that the item parameters are constant over time. That assumption is what makes the exercise possible and it is also its weakest link: THE changed its methodology substantially in 2011 and again in 2024, QS added an employment-outcomes indicator in 2022, and the model absorbs those as changes in institutions rather than in instruments. A fuller treatment would let αⱼ and βⱼ shift at known methodology breaks.

The latent variable is not "quality." It is whatever the rankings jointly measure, which is heavily weighted toward research output, citation impact, Nobel prizes and reputation surveys, and which barely registers teaching, access, or regional service. Reading θ as institutional merit imports every criticism ever made of the underlying rankings. What θ does honestly measure is consensus standing in the global ranking industry — a real and consequential thing, but a narrower one.

Coverage is unbalanced and the gaps are not random. Six of ten systems contribute one to four editions. Five of the ten contribute only to 2025. The dynamic estimates for 2003–2010 rest almost entirely on ARWU and NTU; the estimates for 2023–2026 rest on THE and QS. Everything in between is better identified than either end.

Censoring assumes an eligibility frame. An institution contributes a censored observation to a system-edition only if that system lists it in some other edition. That is a defensible frame but it is a modelling choice, and it means the model never asks why a system ignores an institution entirely.

Entity resolution is good, not perfect. 1073 unresolved candidate pairs remain. For an institution caught in one of those splits, the trajectory will show a spurious break.


7. Files

data/ rankings_panel_long.csv every listing, harmonised: system, year, institution, rank, score crosswalk.csv raw name -> institution id, with all variants edition_summary.csv one row per system-edition: length, censoring cutoff entity_review_candidates.csv 1073 possible-but-unconfirmed same-entity pairs SOURCES.txt repository and commit for every input file estimates/ latent_scores.csv theta posterior mean, sd, 2.5/5/50/95/97.5 percentiles, within-year standardised score, and rank, for every institution-year item_parameters.csv alpha, beta, sigma, reliability per ranking system validation_*.csv edition recovery and pairwise agreement sensitivity_loo.csv leave-one-system-out refits code/ 01_ingest.py read every source into one long file 02_harmonize.py entity resolution 02b_entity_review.py flag unresolved same-entity candidates 03_build_model_data.py edition->reference year, quantile transform, censoring 04_gibbs.py the sampler 05_diagnostics.py convergence, item parameters, validation, estimates 06_figures.py static figures 07_dashboard.py the interactive dashboard 08_sensitivity.py leave-one-system-out 09_memo.py this document figures/ six PNGs university_rankings_dashboard.html self-contained interactive dashboard diagnostics.txt full diagnostic output harmonization_report.txt entity-resolution log with spot checks

Rerun end to end with python3 01_ingest.py && python3 02_harmonize.py && python3 02b_entity_review.py && python3 03_build_model_data.py && python3 04_gibbs.py 15000 && python3 05_diagnostics.py && python3 06_figures.py && python3 07_dashboard.py. Requires pandas, numpy, scipy, rapidfuzz, unidecode, arviz, matplotlib, pyreadr, and the raw files under data/raw/.