Limitations & sources

What the index can and cannot support

An index is only as honest as the limits it states, and stating them once, plainly, is more useful than burying them in a wall of caveats. The points below are the ones a reader needs in order to use the index well and avoid over-reading it. They are not boilerplate: several of them constrain what the current figures can be used for.

Limitations

It measures conditions, not cases.

The index measures the structural conditions associated with forced labor, not its prevalence. A high score means the enabling conditions are strongly present, rather than that a particular number of people are exploited. A low score does not certify a country as free of forced labor.

It reads origin-side structural risk, and under-reads destination and sponsorship systems.

This is the most important thing to understand before reading any single country’s rank. The index is assembled from cross-national indicators that largely describe a country’s resident population and its structural conditions. The mechanisms that drive forced labor in destination economies (kafala-style sponsorship that ties a worker’s legal status to one employer, recruitment-fee debt, and migration brokerage) are named in the framework but are not yet sourced at country scale, and the proxies that are available describe citizens rather than the migrant workforce who are most exposed. As a direct result, several wealthy migrant-destination states score in the lower-risk range: the United Arab Emirates ranks 147th of 184 scored countries, and the other Gulf states sit alongside it, despite well-documented forced-labor risk in their labor systems. Read a low score for a known destination country as “this index does not yet capture that pathway,” not as a clean bill of health. The current index reads origin-side structural vulnerability; it does not yet read destination-side sponsorship capture, and that gap is the next priority for the model rather than a settled feature of the world. Two safeguards make this blind spot visible at the point of use: the eight documented sponsorship-system destinations carry an explicit caveat banner on their country profile and intervention pages, and where Exploitation rests on fewer scored domains the country carries the lower-confidence badge with wider published rank bands — the uncertainty model widens, symmetrically and honestly, rather than guessing which direction the missing data would point.

Read it as bands, not a leaderboard.

A rank-stability analysis that perturbs the modelling choices (weights, the recruitment/exploitation balance, noise on the inputs) finds that countries near the top and the bottom of the index barely move, but positions in the middle of the table sit inside wide intervals: a small data or modelling change can shift a mid-ranked country many places. The honest reading is in tiers: which broad band a country falls in (higher / moderate / lower structural risk) is informative; its exact ordinal rank in the middle of the table is not. Do not treat a difference of a few places mid-table as meaningful.

It is correlated with weak governance, by design.

Risk concentrates where institutions are weak. That is expected, because weak governance is one of the most established structural drivers of forced labor, and the index reports the relationship rather than engineering it away. On the build shown here, the composite correlates with a standard rule-of-law measure at r ≈ 0.79. It is not reducible to governance: roughly a third of the variation in the score is not explained by it, and a pre-registered test confirmed the index still carries a forced-labor-specific signal (the prevalence of child labor) once governance is held constant. It is governance-associated without being governance with a new label. The index also actively de-biases against governance: signals that are empirically too correlated with rule of law are down-weighted rather than entered at full strength. That treatment was stress-tested across every reasonable specification (de-biasing one domain or all of them, and measuring the governance correlation against the World Bank rule-of-law index, the V-Dem rule-of-law index, or a blend of the two), and the index barely moves (rank agreement Kendall τ ≥ 0.96; the governance share stays near 63 percent). The governance handling is not a fragile knob: the result does not depend on which governance measure or scope you choose (see How it’s built).

Eleven countries are not scored, and they are not a random eleven.

The universe is roughly 195 countries; 184 are scored. Where a country’s data are too thin to rank it fairly, the build says so rather than assigning a misleading number. The eleven that drop out are all micro-states and small-island states (Andorra, Dominica, Micronesia, St Kitts & Nevis, Liechtenstein, Monaco, Marshall Islands, Nauru, San Marino, Tuvalu, Vatican City): the same kinds of low-administrative-capacity states for which the labor and governance series are not collected. Because missing data are concentrated in low-capacity states rather than scattered at random, thin coverage tends to pull a score down rather than up; a very low score for a data-sparse country therefore deserves extra caution, not confidence.

Two exploitation domains lean on weak or proxy signals.

Within the exploitation phase, one domain (foreclosed exit) has no direct measure of its core mechanism and is currently carried by proxies for labor-enforcement capacity, and another leans partly on a de-facto governance proxy. So the exploitation half of the score rests more heavily than is ideal on a single well-measured domain. The affected domains are flagged in the data rather than smoothed over, and a reader can see on each country profile which domains were low-confidence or not scored.

Reported figures can be suppressed — and the index does not pretend to correct for that.

Some of the signals the index relies on are themselves recorded by the states being measured, and a state that controls its own records can report conditions as better than they are. The index’s handling is deliberately conservative: a suppressed or absent figure is dropped, never recorded as zero, so opacity can never lower a country’s score — it can only thin the evidence beneath it, and countries whose composite rests on a structurally reduced evidence base are flagged as lower-confidence. What the index does not do is estimate what a suppressing state’s data would have shown: there is no defensible way to test such an adjustment, so none is made. Treat an opaque state’s score as resting on thinner evidence, not as an upper or lower bound.

One recruitment-side indicator runs on thin coverage.

The child-labor prevalence indicator (World Bank/ILO survey series) exists for only 92 of 195 countries, below the index’s 50 percent coverage floor, and the surveys behind it are dated (2005–2016). Consistent with the no-imputation rule, the gap is disclosed rather than filled: the age/childhood domain still scores where enough companion signals exist, but it carries a low-confidence flag, and the indicator is a priority for replacement with modeled estimates in future versions.

The subnational layer covers part of the world.

The regional maps depend on comparable household-survey data, derived from IPUMS-International microdata for 97 countries; the shipped map carries 1,732 admin-1 units across 92 countries, of which roughly 1,450 units in 79 countries pass reliability filtering and carry a risk score. Where it does not exist, the country score remains the best available read. The within-country detail is real but modest for the labor signals at this resolution: it isolates specific high-risk corridors rather than rewriting the national picture wholesale.

There is no clean external prevalence benchmark.

The index was tested against an external estimate of how many people are actually in forced labor (the Walk Free Global Slavery Index 2023 prevalence estimates) and, after netting out governance on both sides, no significant association was found. This is reported plainly, but the benchmark is itself heavily entangled with governance, so a null is uninformative rather than disconfirming. There is no clean, governance-independent measure of forced-labor prevalence to validate against. That absence is the gap a structural index exists to fill, and it is also why the index should be read as a map of conditions rather than a scoreboard of outcomes.

Sources

The index is assembled from established cross-national datasets covering governance, labor, financial access, migration and displacement, gender and age structure, and disaster exposure, among others. Each indicator records its source, vintage, and coverage, so a country’s score can be traced to the data behind it.

The full, per-indicator source list (with the dataset, series, vintage, coverage, licence, and citation for each) is published in the repository as docs/data-provenance.md, and the machine-readable register that drives it is at config/data_register.csv. Required attributions for every bundled source, and a record of the two sources (EM-DAT and IPUMS-International) whose terms restrict redistribution, are kept per-source in docs/data-provenance.md.

For how these sources are combined into a score, see How it’s built.

This index measures the structural conditions associated with forced labor, not its prevalence. It is a tool for identifying where risk is concentrated and why: read alongside, not in place of, on-the-ground knowledge.