What governance indicators' own error bars reveal about what we know and don't know
The World Bank's Worldwide Governance Indicators assign each of 206 countries a score on six dimensions of governance. These scores are used for aid allocation, risk assessment, and academic research. This analysis takes the WGI's own confidence intervals seriously — and finds that they undermine most of the apparent precision the scores are used for.
India is ranked 70th out of 206 countries in Government Effectiveness (score: 59.6). But its 90% confidence interval is [52.7–66.5], meaning India could plausibly rank anywhere from 31st to 132nd — a range of 100 positions. This is not unusual. For any two countries within about 15 points of each other, the data cannot reliably distinguish which has better governance.
| Country | Score | 90% CI | Width | Sources |
|---|
Among these 15 countries, 75% of pairwise comparisons have overlapping confidence intervals. That means for three-quarters of all possible "Country A has better governance than Country B" statements, the data cannot confirm the claim.
Precision is not uniform. The correlation between number of data sources and confidence interval width is r = −0.916 across all 206 countries. Countries covered by more sources get narrower intervals. Nauru, with 2 sources, has a CI width of 27.6 points and could shift ±75 rank positions. These are not measurements — they are rough guesses that appear in the same table as well-measured economies.
Over 17 years (2006–2023), across 10 countries and 4 governance dimensions, only 7 of 40 changes are statistically detectable — meaning the confidence intervals for the first and last year don't overlap. That's 17.5%.
| Country | 2006 | 2006 CI | 2023 | 2023 CI | Change | Detectable? |
|---|
India's Government Effectiveness score rose by 10.7 points over 17 years. This looks like meaningful improvement. But the 2006 interval [42.0–55.8] overlaps with the 2023 interval [52.7–66.5]. The data cannot confirm that governance actually improved.
Year-to-year changes average 1.0–1.7 points. Confidence intervals average 13.0 points wide. Annual "movements" are roughly one-eighth the width of measurement uncertainty. Reports that governance "improved" or "declined" by a point or two are reporting noise as signal.
The six WGI dimensions can tell opposite stories about the same country.
| Country | Corruption | Effectiveness | Rule of Law | Voice | Range |
|---|
Vietnam ranks 3rd in Control of Corruption but 8th in Voice and Accountability. Rwanda ranks 1st in both Corruption and Rule of Law but 5th in Voice. India ranks 1st in Government Effectiveness but 5th in Corruption. A policy brief citing a country's "strong governance" is choosing which dimension to foreground. The choice is the finding.
The WGI has a measurement layer — a set of methodological choices, invisible to most users, that shapes the apparent answer:
The aggregation model weights 35 data sources using an Unobserved Components Model. Sources that correlate more with others receive higher weights. The composite reflects consensus among sources, not necessarily ground truth.
Source coverage varies. Ghana is covered by 15 sources for Rule of Law; Guatemala by 12. The same "score" means different things depending on how much data underpins it.
Confidence intervals exist but are rarely reported. When a governance score is cited in a policy document or news article, the CI almost never accompanies it. The number travels without its uncertainty.
Temporal comparisons assume stability in the instrument. Sources change over time — new surveys added, old ones discontinued. A change in score may reflect a change in governance, a change in sources, or both.
Dimensions are presented separately but used interchangeably. "Governance" is treated as a single concept, but the six dimensions measure different things and rank countries differently.
The WGI are not wrong. They represent a serious, methodologically sophisticated effort to measure something genuinely hard to measure. The confidence intervals are published precisely because the creators understand the uncertainty.
The problem is downstream. Governance scores are used for decisions that require more precision than the data provides. Aid allocation treats small differences as meaningful. Rankings imply an ordering the data can't support. Trend analysis treats year-to-year changes as signal. Dimension selection determines the narrative.
This is the same measurement layer documented in the IATI analysis series: the apparent answer is shaped as much by the measurement methodology as by the underlying reality. The difference is that IATI's layer is about data completeness and reporting expansion, while WGI's layer is about aggregation uncertainty and dimension selection. The pattern is the same.
All data retrieved from the World Bank API (source 3 = WGI). Composite scores, 90% confidence intervals, standard errors, and source counts fetched for six dimensions across 15 focus countries and all 206 economies. Temporal analysis covers 2006–2023. Rank instability calculated by counting positions each country's rank could shift within its 90% CI. Pairwise CI overlap calculated for all country pairs.