Methodology
Version 0.3.0
The composite
The Habitance Score is a weighted geometric mean of eight pillar scores, each normalised to 0–100 by percentile rank within its geographic level. Geometric, not arithmetic, so a failing pillar cannot be papered over by strength elsewhere — that is the survival framing, encoded. Pillars are clamped to [5, 100] before the log so a zero cannot send the composite to negative infinity.
FCS = 100 · exp( Σᵢ wᵢ · ln( clamp(pᵢ, 5, 100) / 100 ) ) if min(p₁…p₈) < 20 → FCS = min(FCS, 40) + a named red flag
The composite is computed in the browser from the pillar scores in the map tiles, not baked into them — which is what lets the horizon, and later a custom weight vector, re-colour the whole globe with no network call.
Red flags: why some scores stop at 40
When any pillar falls below 20 out of 100, the composite is capped at 40 and the failing pillar is named — on the map card, on the place page and on the share card. The geometric mean already punishes a weak pillar; the cap exists because below a certain point the arithmetic stops being the point. A city with excellent hospitals, fast internet and no water is not a mid-scoring city.
A red flag is a statement about one measured pillar, not a verdict on a place — and where our coverage of that pillar is thin we say so beside the flag, because a warning drawn from little evidence is a reason to look closer, not a reason to rule somewhere out. The rule is deliberately blunt and deliberately visible: you can always see the weighted mean before the cap, and which pillar triggered it.
Why some places are scored but not ranked
A place needs at least 40% of the default pillar weight actually measured before it appears in a published ranking. Below that, a position in a league table says more about which indicators reached that place than about the place itself: a unit carrying three pillars out of eight can land in a top 25 because the three it has are its strong ones.
Those places keep everything else — a score, a page, a pin on the globe, a search result — with their confidence figure and their coverage chips saying exactly how thin the evidence is. We would rather show you a number with its limits attached than quietly delete a country from the map because it is hard to measure. The floor applies only to ordered lists, and it lifts on its own as coverage fills in.
Pillars and default weights
| Pillar | Weight | Projectable | What is inside |
|---|---|---|---|
| Climate Resilience | 25% | 2035 · 2050 | Extreme heat, wet-bulb days, drought, flood, wildfire and cyclone exposure. |
| Water Security | 18% | 2035 · 2050 | Baseline and projected water stress, groundwater trend, supply reliability. |
| Stability | 18% | now | Institutional quality, rule of law, and macroeconomic footing. |
| Health & Longevity | 10% | 2050 (UN medium) | Life expectancy, care capacity and air quality. |
| Infrastructure | 10% | now | Power reliability, connectivity, logistics and digital government. |
| Safety | 6% | now | Experienced violence, conflict proximity, and seismic/volcanic hazard. |
| Community & Quality of Life | 8% | now | Social support, education, language access and nature access. |
| Food & Land | 5% | 2050 | Growing conditions, arable land per person, and import dependence. |
A place’s 2050 rank can sit almost exactly where its rank today sits while the physics underneath it moves a great deal. Only a handful of indicators are projected — the climate family, crop yield, and one demographic ratio — and each is normalised by percentile rank within its level at each horizon — so a warming that reaches nearly every place in a cohort moves the median and leaves most ranks where they were. That is a property of rank normalisation, not evidence that nothing changes. Every place page therefore prints the measured quantities behind the climate and water pillars beside the 2050 figure — warmest-month high, degree-months above 30 °C, rainfall variability, coastal-flood exposure, water stress — each against today and against the median of its level. Methodology 0.3.0 addresses the rank itself with absolute anchoring of the physical pillars; the changes register carries the entry.
The Access axis
Access is the second axis, separate from the resilience score: affordability, the legal right to move and to own, and fit. It is an arithmetic mean, not a geometric one — access failures substitute for each other in a way a dry aquifer does not, and a place that is expensive but open is a different case from one that is cheap and closed.
Coverage is uneven, and unevenly in the half that matters most. Affordability and fit are staged for effectively every place. The two indicators that decide whether you can actually go — foreign-ownership rules and visa pathways — come from a hand-compiled table with a citation per row, and reach 8 countries so far: 949 of 7,368 scored places. Everywhere else the composite renormalises over the pillars the place carries, so the Access number — and the quadrant built on it — is about cost and fit alone. The verdict chip says so wherever it appears. A place with no staged Access indicator at all is shown as unscored rather than guessed at, and its quadrant stays open. Read every quadrant as provisional until this section says otherwise.
The pillar handbook
Each pillar above has a chapter of its own: eight short chapters covering what the pillar measures, the indicators and named sources behind it, what it cannot tell you, and two or three real places where it is instructive to read. The chapters read in sequence and recommend nothing.
How much these choices matter
A score is a summary, and every summary rests on choices. These are ours, and this is how much they matter. Every place we score — 7,370 of them — was re-scored under deliberate changes to the model’s own settings: each pillar weight pushed a fifth up and a fifth down in turn, and each of the fifteen heaviest datasets deleted outright.
- 95.3%
- of places stay within 5% of where they were in their own list, under any single weight change
- 0.9612
- lowest rank correlation with the published order after deleting any one of the 15 heaviest datasets
- 297
- places sit on the edge of the 40-point red-flag cap — 4.0% of the world and most of its unstable ranks
The weights barely matter
Move any one of the eight weights by a fifth in either direction and 95.3% of places on Earth stay within five percent of where they were in their own list; 80.0% stay within two percent. The top of every list is almost perfectly still — 203 of the 250 countries move five places or fewer, Iceland does not move at all, Switzerland moves one. If a different reasonable person had chosen different reasonable weights, they would have published very nearly the same rankings.
Nor does any one dataset, taken one at a time
Delete the single heaviest indicator in the model and the world’s order is 99.3% unchanged. Delete any of the 15 heaviest and the order stays at least 96% intact — fourteen of the fifteen leave it above 99%. The floor, 0.9612, is the armed-conflict series, which carries most of the safety pillar’s real texture on its own; its worst median displacement is 58 places out of 7,370. No place’s standing here rests on one file from one agency — but that is the file we watch most closely.
Two things do matter
The first is how crowded a place’s neighbourhood is. Most scores fall in a narrow band, so hundreds of places sit within a point of each other. Crowding predicts a moving rank better than anything else we measured — better than coverage, better than confidence. When a rank moves under a weight change it usually means that place was in a tie, not that the model is unsure about it. Ranks near each other are not meaningfully different: the score, the confidence figure and the pillar bars are the parts to read, and the rank is a rough sort.
The second is the red-flag cap, and it is the largest source of instability in the model — larger than the weights this section opened with. A place with one catastrophic pillar is capped at 40 however good the rest of it is, so 347 cities and 357 regions sit at exactly 40, all tied. 297 places sit on the edge of that line — either narrowly above it, or held at exactly 40 by a fraction of a point. Those 297 are 4.0% of the world and 79% of its 250 most volatile ranks: their rank moves a median of 386 places under a single weight change, against 36 for everyone else. Set them aside and 98.5% of the world moves less than five percent of its cohort. For a capped place the number is set by a rule rather than by a measurement, and its position in a ranking should be read as “capped”, not as “2,202th”. The red flag on the card says so.
Worth separating two numbers that get confused: 3,010 places (40.8%) carry a red flag, but only 756 (10.3%) have a score the cap actually decided. The other 2,254 would have scored at or below 40 on their own.
| Place | Level | Score | Rank | Worst move |
|---|---|---|---|---|
| Iceland | Country | 79.25 | 1 | 0 places |
| Ottawa | City | 80.40 | 6 | 0 places |
| Switzerland | Country | 75.98 | 3 | 1 place |
| Caracas | City | 37.89 | 2,726 | 529 places |
| Bangkok | City | 39.11 | 2,627 | 433 places |
| Johannesburg | City | 40.00 | 2,202 | 388 places |
Caracas, Bangkok and Johannesburg are not places the model is confused about. All three sit within a couple of points of the 40-point cap, and a point or two of score there buys hundreds of places of rank because it crosses the tie block. Their rank is the position of a threshold, not an opinion about the city.
Where we know least, we say so — and we do not rank
The confidence figure earns its place at the extremes rather than in the middle: at the 90th percentile, the 16 least-confident countries have ranks that wander 26.81% of the country list against 7.20% for the rest. Across the rest of the distribution confidence separates places only weakly — and at this version, some low-confidence chips mark countries whose figures are measured region by region rather than poorly known at all, so the chip alone no longer predicts a volatile rank. Saying so plainly matters more than the correlation does. The coverage floor above is carrying more of this work: the most volatile units in the model are the ones already labelled “insufficient data to rank” rather than listed.
Full report, with the per-indicator and per-level tables and the four things it does not claim: docs/SENSITIVITY_REPORT.md (2026-09-01, methodology 0.2.0, seed 20260828, 500 draws, reproducible with uv run marts sensitivity --full). It is an internal-consistency study: showing that a ranking is insensitive to its weights says nothing about whether the right things are being weighted, and dropping an indicator is not the same test as an indicator being wrong.
What we will not claim
- Only an indicator that declares a projection carries a 2035 or 2050 value: the climate family under an emissions scenario, and the demographic family under the UN's medium variant. Stability, infrastructure, safety and community hold their current value at every horizon — projecting a government to 2050 is fiction, and the model refuses to. A “2050 score” is projected climate, water, crop yield and the UN's medium population variant, against today's stability, infrastructure, safety and community, and is labelled that way everywhere. It is not a forecast of 2050 society.
- Confidence is displayed next to every score and is never hidden.
- No pay-to-rank and no sponsored scores. Affiliate links are labelled.
Data freshness
Every source in the registry, its licence tier, and how old the data actually is. Built from sources.yaml and the ingest manifests, so it cannot drift from what the pipeline is really reading.
| Source | Version | As of | Days | Fetched | Cadence | Tier |
|---|---|---|---|---|---|---|
| BIS Effective Exchange Rate indices (WS_EER)bis_eer · not in the current composite | 2026-08-26 | — | — | Aug 30, 2026 | monthly | Free tier only |
| BIS Selected Residential Property Prices (WS_SPP)bis_property_prices · not in the current composite | 2026-08-26 | — | — | Aug 30, 2026 | quarterly | Free tier only |
| CEMS fire danger indices, historical (GEFF forced by ERA5), system 4.1 — EWDScopernicus_cems_fire_historical · 1 indicator in use | v4_1-consolidated_dataset | Jan 1, 2014 | 4636 | Sep 9, 2026 | daily (extent 1940-01-03 -> 2026-08-31; cads:update_frequency "Daily") | Pro / API |
| Copernicus GLO-30 Digital Elevation Modelcopernicus_glo30 · not in the current composite | 2021-1 | — | — | Aug 30, 2026 | static (versioned) | Pro / API |
| Desalination plant capacity by country (Habitance compilation)desalination · 1 indicator in use | 2026-08-30 | Aug 28, 2026 | 14 | — | annual; plus on commissioning of any `under_construction` plant | Pro / API |
| DOSE — MCC-PIK Database Of Subnational Economic output (V2.9)dose · not in the current composite | v2.9 | — | — | Aug 31, 2026 | irregular (V2.9 published 2024-09-17; concept DOI 10.5281/zenodo.7573249 tracks latest) | Pro / API |
| EF English Proficiency Indexef_epi · not in the current composite | — | — | — | — | annual | Free tier onlyunverified |
| Ember Yearly Electricity Dataember_electricity · 1 indicator in use | 2026-08-10 | Dec 31, 2025 | 254 | Aug 30, 2026 | annual | Pro / API |
| EOG VIIRS Nighttime Lights annual composites (VNL v2.2, average_masked)viirs_vnl · not in the current composite | v2.2 | — | — | Sep 2, 2026 | annual | Pro / API |
| Eurostat regional statistics (NUTS2/NUTS3)eurostat_regional · not in the current composite | 2026-08-31 | — | — | Aug 31, 2026 | annual (per-dataset update stamps in the API metadata) | Pro / API |
| FAO AQUASTATfao_aquastat · not in the current composite | — | — | — | — | annual | Free tier onlyunverified |
| FAOSTAT bulk CSV — food security, SDG indicators, crop productionfaostat · 3 indicators in use | 2026-07-21 | Dec 31, 2025 | 254 | Aug 30, 2026 | continuous per domain (each carries its own DateUpdate) | Pro / API |
| Foreign property-ownership rules and residence-visa pathways (Habitance compilation)foreign_ownership_visa · 2 indicators in use | 2026-08-30 | Aug 28, 2026 | 14 | — | semi-annual (docs/03_DATA_SOURCES.md §9); ad hoc on any reform | Pro / API |
| Fragile States Index (Fund for Peace)fsi · not in the current composite | 2023 | — | — | Aug 29, 2026 | annual | Free tier only |
| Frankfurter exchange-rate API (central-bank reference rates)frankfurter_fx · 1 indicator in use | 2026-08-29 | Aug 29, 2026 | 13 | Aug 30, 2026 | daily (use monthly downsample) | Pro / API |
| Freedom House — Freedom in the Worldfreedom_house · not in the current composite | FIW2025 | — | — | Aug 30, 2026 | annual (February/March) | Free tier only |
| geoBoundaries (gbOpen) administrative boundariesgeoboundaries · not in the current composite | gbOpen-9469f09 | — | — | Aug 29, 2026 | rolling releases | Free tier only |
| GeoNames gazetteer (cities15000)geonames · not in the current composite | 2026-08-29 | — | — | Aug 29, 2026 | daily dump | Pro / API |
| Germanwatch Climate Risk Indexgermanwatch_cri · not in the current composite | 2026 | — | — | — | annual | Free tier only |
| GHM_drought — individual-model SPEI from 16 CMIP6 models, 1961–2100 (Data product 2)ghm_drought_spei_cmip6 · not in the current composite | zenodo-21491403 | — | — | — | static (Zenodo record, published 2026-07-22) | Pro / API |
| GHSL Global Human Settlement Population Grid (JRC)ghsl_pop · not in the current composite | R2023A-E2025 | — | — | Aug 29, 2026 | periodic (R2023A is the current global release) | Pro / API |
| Global Roads Inventory Project v4 (GRIP4) — global road density rasters, PBL / GLOBIOgrip4 · 1 indicator in use | v4-2018-04 | Apr 24, 2018 | 3062 | Sep 9, 2026 | static (rasters dated 2018-04-24; road sources 2000–2015; no upstream versioning) | Pro / API |
| Heritage Foundation Index of Economic Freedomheritage_efi · not in the current composite | 2026 | — | — | Aug 30, 2026 | annual | Free tier only |
| HydroSHEDS v1 (HydroRIVERS / HydroBASINS)hydrosheds · not in the current composite | v1.0 | — | — | Sep 1, 2026 | static (v1.0 current; v2 in progress upstream) | Free tier only |
| IIASA SSP Scenario Database — basic drivers release 3.2 (population, GDP by SSP, 2025–2100)iiasa_ssp · not in the current composite | release-3.2-2026-01-27 | — | — | Sep 2, 2026 | static (release 3.2, May 2025; file Last-Modified 2026-01-27) | Free tier only |
| IMF Global Housing Watchimf_housing_watch · not in the current composite | — | — | — | — | quarterly | Free tier onlyunverified |
| IMF International Reserves and Foreign Currency Liquidity (IRFCL)imf_irfcl · not in the current composite | 12.0.0 | — | — | Aug 30, 2026 | monthly | Free tier only |
| IMF World Economic Outlook Databaseimf_weo · not in the current composite | 9.0.0-2026-04 | — | — | Aug 30, 2026 | 2x/yr (April and October) | Free tier only |
| INFORM Risk Index (EC JRC / DRMKC)inform_risk · not in the current composite | — | — | — | — | 2x/yr (September and March) | Pro / APIunverified |
| Institute for Economics & Peace — Global Peace Index and Global Terrorism Indexgpi_gti · not in the current composite | — | — | — | — | annual (GPI June, GTI March) | Free tier onlyunverified |
| International Comparison Program 2021 results — price level indices (World Bank API database 90)wb_icp_2021 · 2 indicators in use | 2024-08-04 | Dec 31, 2021 | 1715 | Sep 9, 2026 | benchmark cycle (2011, 2017, 2021; the API's lastupdated is 2024-08-04) | Pro / API |
| International Property Rights Index (Property Rights Alliance)ipri · not in the current composite | 2025 | — | — | Aug 30, 2026 | annual | Free tier only |
| IPCC AR6 WG1 Chapter 9 Table 9.9 — global mean sea level projectionsipcc_ar6_slr · not in the current composite | wg1-ch9-table-9.9 | — | — | Aug 29, 2026 | static (assessment report; next revision is AR7) | Pro / API |
| ISIMIP3b Climate Extremes Indicators (SPARCCLE2025 derived input data), v1.1isimip3b_climate_extremes · not in the current composite | v1.1-20260206 | — | — | — | static (versioned; v1.1 = 20260206) | Pro / API |
| ISIMIP3b GGCMI phase 3 crop-yield simulations (OutputData/agriculture) + landuse-15crops 2015soc inputisimip3b_ggcmi · not in the current composite | ISIMIP3b-OutputData-agriculture-2026-09-02 | — | — | — | static (ISIMIP3b round; per-dataset DOIs) | Pro / API |
| NASA Global Landslide Susceptibility Map (Stanley & Kirschbaum 2017; 2023 GeoTIFF)nasa_landslide_susceptibility · not in the current composite | 2023-02-27 | — | — | Sep 2, 2026 | static (file Last-Modified 2023-02-27) | Free tier only |
| NASA NEX-GDDP-CMIP6 (COG sub-products)nex_gddp_cmip6 · 4 indicators in use | cog-2022-10 | Jan 1, 2012 | 5367 | Aug 29, 2026 | static (versioned; registry states "No future updates planned") | Pro / API |
| NASA NEX-GDDP-CMIP6 (daily per-model NetCDF archive)nex_gddp_cmip6_daily · not in the current composite | v2.0 | — | — | — | static (versioned per file — base, _v1.1, _v2.0; the lane takes the newest per file) | Pro / API |
| Natural Earth 10m cultural vectors (admin-0, admin-1, coastline)natural_earth · 1 indicator in use | 10m-v5.1.1 | Jan 1, 2022 | 1714 | Sep 1, 2026 | static (versioned) | Pro / API |
| Natural Earth 10m physical vectors + 50m pre-shaded relief rastersnatural_earth_physical · not in the current composite | 10m-v5.1.1 / raster-50m-v2.0.0 | — | — | — | static (versioned; vectors v4.0.0–v5.1.1 per theme, rasters v2.0.0) | Pro / API |
| ND-GAIN Country Index (Notre Dame Global Adaptation Initiative)nd_gain · not in the current composite | 2026 | — | — | Aug 29, 2026 | annual | Free tier only |
| NOAA ETOPO 2022 global relief (60 arc-second)etopo · 1 indicator in use | 2022-v1-60s | Jan 1, 2022 | 1714 | Aug 29, 2026 | static (versioned) | Pro / API |
| NOAA IBTrACS — International Best Track Archive for Climate Stewardshipibtracs · not in the current composite | v04r01 | — | — | Aug 30, 2026 | ongoing | Pro / API |
| OECD Cities / Functional Urban Areas (SDMX 2.1) — DSD_FUA_* familyoecd_fua · not in the current composite | 2026-09-01 | — | — | Sep 2, 2026 | annual | Pro / API |
| OECD Data Explorer (SDMX 2.1) — analytical house prices + regional TL2 indicatorsoecd_data_explorer · 3 indicators in use | 2026-08-29 | Jun 30, 2026 | 73 | Aug 30, 2026 | quarterly (prices) / annual (regional) | Pro / API |
| OECD Economic Outlook 117 long-term scenarios (to 2060)oecd_eo_ltb · not in the current composite | EO117 | — | — | Sep 2, 2026 | with each long-term Economic Outlook update (EO117 = June 2025) | Pro / API |
| OECD Regional Statistics (SDMX 2.1) — TL2 social/health/economy/education waveoecd_regional · 2 indicators in use | 2026-08-31 | Dec 31, 2025 | 254 | Sep 1, 2026 | annual | Pro / API |
| OurAirports global airport databaseourairports · 1 indicator in use | 2026-08-29 | Dec 31, 2026 | 0 | Aug 30, 2026 | nightly | Pro / API |
| Protomaps planet basemap (OpenStreetMap-derived) + basemaps-assets glyphsprotomaps · not in the current composite | 20260830 | — | — | Aug 30, 2026 | daily planet builds (we pin one build key and re-pin deliberately) | Pro / API |
| RESOLVE Ecoregions 2017 (846 terrestrial ecoregions, 14 biomes)resolve_ecoregions_2017 · not in the current composite | 2017 | — | — | Sep 1, 2026 | static (2017 one-time release; file last-modified 2019-05-07) | Pro / API |
| SatPM2.5 / van Donkelaar V6.GL.03 surface PM2.5 (WashU ACAG)pm25_satpm · 1 indicator in use | V6GL03 | Dec 31, 2024 | 619 | Aug 30, 2026 | annual | Pro / API |
| Smithsonian Global Volcanism Program — Volcanoes of the Worldgvp_volcanoes · 1 indicator in use | 2026-08-29 | Dec 31, 2026 | 0 | Aug 30, 2026 | ongoing | Pro / API |
| STORM tropical cyclone wind speed return periods — present climate (v4) + STORM-C climate change (v2)storm_tc_wind · 1 indicator in use | present-v4-2023-06-22_future-v2-2022 | Jun 22, 2023 | 1177 | — | static | Pro / API |
| UN E-Government Development Index (UN DESA)un_egdi · not in the current composite | — | — | — | — | 2-yearly | Free tier onlyunverified |
| UN World Population Prospects 2024un_wpp · 2 indicators in use | WPP2024 | Jul 1, 2023 | 1168 | Aug 29, 2026 | 2-yearly | Pro / API |
| UNESCO Institute for Statistics (UIS) public APIunesco_uis · not in the current composite | 20260507-91260335 | — | — | Aug 30, 2026 | 2x/yr data releases | Free tier only |
| UNODC dataUNODC — intentional homicide and crime statisticsunodc · not in the current composite | — | — | — | — | annual | Free tier onlyunverified |
| Uppsala Conflict Data Program — GED, BRD, Non-stateucdp · 1 indicator in use | 26.1 | Dec 31, 2025 | 254 | Aug 30, 2026 | annual (plus monthly candidate releases) | Pro / API |
| US Census Bureau — American Community Survey 5-year data profiles + Gazetteerus_census_acs · not in the current composite | 2024 | — | — | Sep 2, 2026 | annual | Pro / API |
| USGS Design Maps web service (ASCE 7-16 seismic design parameters)usgs_seismic · 1 indicator in use | asce7-16-v4.0.x | Dec 31, 2016 | 3541 | Aug 30, 2026 | static per model version | Pro / API |
| V-Dem (Varieties of Democracy) Country-Year Full+Others v16vdem · not in the current composite | v16 | — | — | Aug 30, 2026 | annual (March) | Free tier only |
| WHO Global Health Observatory (OData API)who_gho · not in the current composite | 2026-08-29 | — | — | Aug 29, 2026 | annual | Free tier only |
| WHO/UNICEF Joint Monitoring Programme (washdata.org) — household estimatesjmp_washdata · not in the current composite | 2026-08-29 | — | — | Aug 30, 2026 | annual | Free tier only |
| Wikidata — the `image` (P18) statement per country / region / city itemwikidata · not in the current composite | 2026-09-10 | — | — | Sep 11, 2026 | live (the SPARQL answers are pinned by fetch date under data/raw/wikidata/<version>/) | Pro / API |
| Wikimedia Commons — `imageinfo` (licence, author, credit, 1280 px thumbnail) per filewikimedia_commons · not in the current composite | 2026-09-10 | — | — | Sep 11, 2026 | live (answers pinned by fetch date; thumbnails staged under data/staged/photos/<version>/) | Pro / API |
| World Bank Enterprise Surveys outage block, as republished in WDI (database 2)wdi_enterprise_surveys · not in the current composite | 2026-07-13 | — | — | Sep 9, 2026 | survey rounds (irregular per country; WDI refresh, lastupdated 2026-07-13) | Pro / API |
| World Bank Human Capital Indexwb_hci · 1 indicator in use | 2020-09-21 | Dec 31, 2020 | 2080 | Aug 30, 2026 | none — frozen | Pro / API |
| World Bank International Comparison Program — PPP conversion factorswb_icp · not in the current composite | 2026-07-13 | — | — | Aug 30, 2026 | periodic (ICP rounds; series refreshed with WDI) | Pro / API |
| World Bank Logistics Performance Indexwb_lpi · 1 indicator in use | 2026-07-13 | Dec 31, 2022 | 1350 | Aug 30, 2026 | every 2-3 years (survey) | Pro / API |
| World Bank WDI — natural-resource rents (oil, gas, coal, total)wb_rents · not in the current composite | 2026-07-13 | — | — | Sep 2, 2026 | continuous (WDI `lastupdated` 2026-07-13) | Pro / API |
| World Bank WDI — water block (AQUASTAT and JMP series, republished)wdi_water · 2 indicators in use | 2026-07-13 | Dec 31, 2024 | 619 | Aug 30, 2026 | continuous (WDI refresh) | Pro / API |
| World Bank World Development Indicators — macro, infrastructure, safety, food and price-level blockwdi_macro · 11 indicators in use | 2026-07-13 | Dec 31, 2025 | 254 | Aug 29, 2026 | continuous (WDI refresh; `lastupdated` stamps each response) | Pro / API |
| World Bank Worldwide Governance Indicatorswgi · 4 indicators in use | 2026-03-18 | Dec 31, 2024 | 619 | Aug 29, 2026 | annual | Pro / API |
| World Happiness Report — data appendixwhr · 1 indicator in use | 2026 | Dec 31, 2025 | 254 | Aug 29, 2026 | annual (March) | Pro / API |
| WRI Aqueduct 4.0 Water Risk Atlasaqueduct_40 · 4 indicators in use | v4.0-2023-08-16 | Jan 1, 2023 | 1349 | Aug 29, 2026 | static (v4.0 released 2023-08-16; no v4.1 published) | Pro / API |
| WRI Aqueduct Floods hazard maps v2aqueduct_floods · 1 indicator in use | v2-2026-03-27 | Mar 27, 2026 | 168 | Aug 30, 2026 | static | Pro / API |
“As of” is the newest observation the source contributes to a live score, and “Days” is its age at the last build (Sep 11, 2026) — not at the moment you are reading. “Fetched” is when the raw file was last pulled, which is a different question: a fresh download of a 2023 release is still 2023 data. A source listed as “not in the current composite” is registered and downloaded but does not yet feed a published indicator. “Free tier only” marks a source whose commercial licence is unconfirmed or restricted; it is displayed and withheld from Pro exports and the API (CLAUDE.md §3 rule 3).
Full document
The methodology as written — docs/METHODOLOGY.md at version 0.3.0, verbatim. The sections above are its summary, and every score on this site was computed under it.
Read the full methodology text
# FUTURECAST — Methodology
```yaml
methodology_version: 0.3.0
status: frozen # opened by the 2026-09-02 [HUMAN] ballot
# (docs/METHODOLOGY_v0.3_BALLOT.md — all 15 approved,
# Eric 2026-09-03: "APPROVE all, keep streaming");
# FROZEN 2026-09-09 (freeze ballot F1–F5 signed), scored
# 2026-09-10, deployed 2026-09-10 — the record is §18.
# 0.2.0 remains frozen and reproducible from
# reference_library/0.2.0/; its record is §17, unedited.
authored_by: "[scoring] agent"
date: 2026-09-03
opened_on: 2026-09-03
frozen_on: 2026-09-09
opened_by: "docs/METHODOLOGY_v0.3_BALLOT.md — [HUMAN], Eric Rehm (approve all 15)"
reference_library: data/staged/reference_library/0.3.0/ # ranked ids only — ballot A1
anchoring_table: packages/scoring/src/futurecast_scoring/methodology/v0_3.yaml → anchoring
# the frozen value → score curves — ballot A1
supersedes: 0.2.0 # minted and immutable; its record is §17, kept unedited
binding_inputs:
- docs/REVIEW.md # overrides brief + architecture (CLAUDE.md §5)
- docs/REVIEW.md addendum 2026-08-29-b # D1-D8, B1-B7; D5 and D8 are implemented here
- docs/REVIEW.md D9, D10 # the manual-table review and the freeze ballot
- docs/REVIEW.md D12 # the rule-7 demotions — see §16
- docs/PHASE_0_DONE.md # the Phase-0 [HUMAN] gate record, folded into §2, §3.6, §13
- docs/01_PRODUCT_BRIEF.md §3
- docs/02_ARCHITECTURE.md §3.3, §4
- docs/DATA_ISSUES.md ISSUE-005 # heat indicators redefined, see §3.3
- docs/METHODOLOGY_v0.2_BALLOT.md # the 0.2.0 ballot — A1–A6, B1–B2, C1–C2, D1–D5
- docs/METHODOLOGY_v0.3_BALLOT.md # the 0.3.0 ballot — A1–A5, B1, C1–C4, D1, E1–E3
implementation: packages/scoring
config: packages/scoring/src/futurecast_scoring/methodology/v0_3.yaml # v0_1, v0_2 frozen
freeze_record: "§15 is 0.1.0's, §16 is 0.1.1's, §17 is 0.2.0's (frozen 2026-09-01); §18 carries 0.3.0's when it freezes"
```
This document is normative. Every formula below is implemented exactly, once, in
`packages/scoring`; the package contains no second definition of anything stated here.
Where this document and the brief disagree, this document (and `REVIEW.md` behind it)
wins. Where this document and the code disagree, that is a bug in the code.
**This document now carries `0.2.0`**, opened on **2026-09-01** by the [HUMAN] ballot
`docs/METHODOLOGY_v0.2_BALLOT.md` — all fifteen decisions approved as recommended, which is
the **[HUMAN]** gate CLAUDE.md §2 and §12 require for an indicator-set change. `0.1.0` (frozen
2026-08-30, REVIEW **D10**) and `0.1.1` (same day, REVIEW **D12**) are minted and immutable:
every number published under `0.1.1` remains reproducible from §15–§16 of this document, from
`packages/scoring/src/futurecast_scoring/methodology/v0_1.yaml` (kept frozen on disk), and from
the reference distributions frozen at `data/staged/reference_library/0.1.1/`; every number
published under `0.1.0` is still reproducible from `0.1.0/`, which was never opened for
writing. `0.2.0`'s config is `v0_2.yaml` and its reference library is rebuilt from scratch
into `0.2.0/` (ballot D2). Changing a weight, a threshold, a formula or the indicator set does
not edit a frozen version — it opens a new one (§12), which is exactly what the 0.2.0 ballot
did.
**Two records, not one.** **§15 is `0.1.0`'s** — no longer a checklist but the record of what
was true at that freeze, **kept unedited** so a reader in a year can see what was known and
what was owed, including the two licence exposures it shipped with. **§16 is `0.1.1`'s**: the
demotions that closed them, what they cost, and what would bring the two indicators back.
---
## 0. What this score is
*(Added 2026-08-29 on REVIEW addendum 2026-08-29-b D8. Everything below §0 leads with
limitations, and a reader who meets the limitations before the product concludes there is no
product. The limitations are right. This section comes first.)*
**The Habitance Score answers one question: will this place still work in thirty years?**
It is a 0–100 number for every country, region and city on Earth, built from the physical and
institutional facts that decide whether somewhere holds up — how hot and how dry it gets, how
much water it has, whether its institutions function, whether it can feed and treat and connect
the people living there. High means a place whose foundations look durable. Low means a place
running down a resource, an institution, or a climate it cannot replace.
Beside it sits a second, independent number, the **Access Score**: can *you* actually get in?
Price, legal right to buy, visa routes. The two never blend, because a place can be resilient and
closed (Norway) or fragile and wide open (a cheap coastal market with a golden visa), and one
number averaging those would describe neither. The quadrant chip — Move/Buy, Wait, Trap,
Avoid — is the compression that makes the pair readable at a glance.
**What it tells you.** How places rank against each other on durability, today and on a 2050
climate horizon. Which specific thing is the weak link — the model names the failing pillar
rather than burying it in an average. How much a place's standing moves once the climate horizon
is applied, and whether that movement is bigger than the model's own uncertainty.
**What it deliberately does not tell you.**
- **It is not a forecast.** Only the climate and water pillars carry 2035/2050 values, and even
those are *scenarios* under a named emissions pathway, not predictions. Nothing else is
projected — a stability score is never labelled 2050 (§8).
- **It does not describe a property.** The finest unit is a city or an admin-1 region. A score
is about a place, not a parcel, and the model refuses to imply otherwise (§14).
- **It is not advice.** Informational only; not financial, legal, or relocation advice.
- **It ranks; it does not measure against an absolute standard.** A score is a percentile against
a frozen peer set, so "80 on climate" means "better than four-fifths of its cohort", not
"safe" (§4.1). This is the single most important limitation and the first thing v0.2 fixes.
- **It scores no culture, religion, ethnicity or politics of the people living somewhere.**
Demographic and cultural context appears on the card as *information*, for the reader's own
judgement, and enters no score at any weight (REVIEW D1).
**Why the caveats are the product, not the fine print.** Anyone can publish a number. What makes
a resilience score worth paying for is that you can check it: every value on the site resolves to
a named source, a named version and a named transform; every score ships with a confidence figure
and a "thin data" flag when a pillar is thinly covered; every place where a quantity we wanted
was not buildable is written down as a limitation with the substitute named (§3.3, §3.4, §3.7).
Competitors in this space publish a number and a marketing page. The caveats below are the
difference, and they are why the rest of this document reads the way it does.
---
## 1. What is being computed
Two independent axes, never blended (REVIEW §1 Q1):
| Axis | Range | Aggregation | Question |
|---|---|---|---|
| **Habitance Score (FCS)** | 0–100 | weighted **geometric** mean of 8 pillars | Will this place hold up? |
| **Access Score** | 0–100 | weighted **arithmetic** mean of 3 pillars | Can I actually get in? |
Both carry a **confidence** in `[0, 1]` that is displayed everywhere the score is
displayed (CLAUDE.md §3.6). Never show a score without its confidence.
### 1.1 Why the aggregations differ (the substitutability asymmetry)
The two axes use different means on purpose, and the difference is the methodological
claim of the product.
**FCS is geometric because resilience failures are not substitutable.** A dry aquifer
is not offset by good schools. There is no amount of Community score that makes water
appear. The geometric mean encodes exactly that: it is dominated by the smallest term,
so a 15 in Water cannot be papered over by 90s elsewhere. This is the survival framing,
in arithmetic.
**Access is arithmetic because access failures *are* substitutable.** A high price is
genuinely offset by an easy visa: money and legal friction trade against each other in
the real decision a buyer makes. A 20 in Affordability plus a 90 in Legal access is a
real, findable situation (an expensive country with a nomad visa), and averaging it to
~50 describes it correctly. Using a geometric mean here would falsely claim that
expensive-but-open is as closed as cheap-but-banned.
State this asymmetry on `/methodology`. It is the honest answer to "why two different
formulas?" and it is a better talking point than either formula alone.
---
## 2. Pillars and default weights
Eight FCS pillars. Default weights per REVIEW §1 Q3 (sum = 100):
| # | Pillar id | Label | Default weight | Horizon-eligible |
|---|---|---|---|---|
| 1 | `climate` | Climate Resilience | **25** | **yes** |
| 2 | `water` | Water Security | **18** | **yes** |
| 3 | `stability` | Stability | **18** | no |
| 4 | `health` | Health & Longevity | **10** | no |
| 5 | `infrastructure` | Infrastructure & Connectivity | **10** | no |
| 6 | `safety` | Safety | **6** | no |
| 7 | `community` | Community & Quality of Life | **8** | no |
| 8 | `food` | Food & Land Resilience | **5** | no |
Weights changed from brief §3.1 on REVIEW's instruction: Water 15 → 18 (least
substitutable physical constraint), Stability 20 → 18 and Safety 7 → 6 (conflict content
de-duplicated between them, §3.1 below), Climate/Health/Infra/Community/Food unchanged.
The Access Score has three pillars:
| Pillar id | Label | Default weight |
|---|---|---|
| `affordability` | Affordability | **45** |
| `legal_access` | Legal access | **35** |
| `fit` | Fit | **20** |
> **Confirmed at the Phase-0 gate, provisionally.** Neither the brief nor REVIEW specified
> Access weights; 45/35/20 was a scoring-agent default (affordability is the binding
> constraint for most users; legal access is binary-ish and decisive when it bites; fit is
> preference, not constraint). Eric confirmed it at the 2026-08-29 demo gate as
> *"for now; tweak with new data"* (`docs/PHASE_0_DONE.md`). It is therefore a **decision,
> not an open question** — and it was taken against an Access axis with one staged indicator,
> so the intent is explicitly to revisit it once `legal_access` and `fit` carry values. That
> revisit is a **minor** bump and a **[HUMAN]** gate (§12); it is on the §15 checklist.
---
## 3. Indicator → pillar assignment
### 3.1 Assignment rules (binding, REVIEW §2)
1. **Every scoring indicator feeds exactly one FCS pillar.** No indicator appears in two
FCS pillars. (An indicator may additionally appear in the Access axis — the two axes
are independent models, not two halves of one model. `english_proficiency` is the
only such reuse in v0.1 — and since 2026-08-30 it is display-only on both axes
(§3.6.4), so v0.1 currently ships with no scored reuse at all.)
2. **Conflict lives in Safety only.** Stability keeps *institutional* measures (WGI, the
economic block; V-Dem until its licence question is settled). Safety keeps *experienced*
violence (homicide, conflict events — terrorism is no longer carried, §3.6.4) plus
**natural hazard (seismic, volcanic), which is always on** — hazard is not "personal
defense" and is never persona-gated.
3. **Composites of our own inputs are display-only.** The test, stated once so the
dispositions follow from a rule rather than from taste:
> **Is this index a weighted combination of quantities that are already in our model?**
> If yes, scoring it counts them twice and imports someone else's black box. It becomes
> `display_only: true`, weight 0, shown on the Place Card as corroboration.
Applied, with the sentence that must accompany each one wherever it is displayed:
| Index | Disposition | The sentence |
|---|---|---|
| **Fragile States Index** | display-only | *Built substantially on the same governance signals as the WGI measures in our Stability pillar, so we show it beside our score rather than inside it.* |
| **Global Food Security Index** | display-only | *A composite over arable land, import dependence and agro-suitability, which are three of our own Food indicators; it corroborates the pillar, it does not add to it.* (REVIEW D3 — excluded by this rule, **not** by licence.) |
| **OECD Regional Well-Being index** | display-only | *A composite over eleven topics — income, health, safety, environment, education — most of which we score directly. Its **components** are used individually; `social_support` is its Community topic.* |
| **ND-GAIN, INFORM** | display-only | Composites over vulnerability and readiness signals we score directly. |
| **Global Peace Index** | ~~display-only~~ → **removed** 2026-08-30 | Rule 3 demoted it; the licence then removed it. See below. |
| **Global Terrorism Index** | ~~scored 0.16~~ → **retired** 2026-08-30 | Rule 3 kept it; the licence still removed it. See below. |
**What happened to the two IEP indices, 2026-08-30 (DATA_ISSUES ISSUE-111).** The Institute
for Economics & Peace sells commercial licences and defines *any* use by an organisation
that is not a non-profit or research institution as commercial use. **Display on a
commercial site is therefore commercial use**, which means the usual remedy — drop the
weight, keep the index on the card as corroboration — was not available. GPI was removed
from the model outright (it had never appeared in a staged or published mart, which is what
made deleting the id permissible at all, CLAUDE.md §4). GTI was retired: the id survives at
weight 0 with empty `sources` and empty `expected_sources`, so nothing can be ingested and
nothing can be rendered. Its 0.16 went to `homicide_rate` and the conflict-event
indicator (§3.6.4; that id was itself deprecated on 2026-08-30 and the weight now sits on
`conflict_events_per_100k` — §3.6.6). **Rule 3's reasoning about GTI was and remains sound — GTI added information the
model did not otherwise carry — and that is precisely why the loss is real rather than
tidy.** Safety now measures organised violence (UCDP) rather than terrorism as such.
*(FSI's display-only disposition was REVIEW §2's explicit ruling; GFSI's was ratified by
REVIEW addendum 2026-08-29-b D3 in Eric's own words; OECD-RWB extends the same rule. It is
on the §15 checklist for explicit sign-off at the freeze anyway. The GPI line that used to
carry that caveat is now moot — the index is gone for a different reason entirely.)*
4. **Personal-defense indicators are persona-gated** (REVIEW §1 Q5). `firearm_legality`
and `distance_to_active_conflict_km` are **excluded from the default weight vector
entirely** and enter the model only under a persona that names them (v0.1: only
`resilience_first`). `conflict_events_per_100k` is *not* gated — it is experienced
violence and stays always-on in Safety.
5. **Persona-gated weights are additive.** Within each pillar the *default-active*
indicator weights sum to exactly 1.0. A gated indicator carries its own weight on top;
enabling it renormalizes the pillar (§5) rather than silently rescaling the defaults.
This is checked by a validator at load time.
6. **No indicator whose source forbids scraping or whose license blocks commercial use
is in the scoring set.** Numbeo QoL is therefore absent (CLAUDE.md §3.4). The license
gate for the sources that *are* present is enforced elsewhere
(`tests/test_license_gate.py`, QA agent).
7. **A scored indicator's sources must all be commercially cleared** — added 2026-08-29,
binding at the `0.1.0` freeze. Rule 6 was written against sources whose licence is
*decided* against us; this rule closes the case rule 6 left open, which is a source that
is merely `review` or `unverified`.
The reason is structural rather than legal. **There is exactly one FCS per (unit, persona,
horizon).** It is the same number on the free tier, in a Pro export, in an API response and
on a share card. So an indicator fed by a source that is not
`license_ok_commercial: true` **and** `status: verified` cannot carry weight: scoring with
it either leaks a restricted source into a paid response, or forks the score into two
different numbers for one place — and a place with two scores has none.
Machine form, **two halves, and the second was added because the first alone was not
enough**:
1. **no scored indicator may carry `tier: free_only`** in `indicators/*.yaml` — the
*labelling* check, `tests/qa_lib.validate_indicator_tiers`;
2. **no indicator fed by an uncleared source may carry `weight_in_pillar > 0`** in
the current version's methodology YAML (`methodology/v0_2.yaml` at `0.2.0`) — the
*weight* check,
`tests/qa_lib.validate_restricted_indicator_weights`, added 2026-08-30 by REVIEW **D12**.
Half 1 was in place for the whole of `0.1.0` and green, and `0.1.0` still shipped two
scored indicators on `review` carriers, because both were labelled correctly and the gate
had no opinion about their weight (ISSUE-132). Labelling a number `free_only` does not
create a second, cheaper score to keep it in; there is only ever the one. Half 2 is what
makes the rule enforceable, and it is fixture-tested against a deliberate violation.
An indicator whose source is restricted becomes `display_only` — it may still appear on the
card, marked, because free-tier *display* is what a restricted licence permits. Restoring
it when its licence clears is an indicator-set change ⇒ **major** bump and a **[HUMAN]**
gate (§12).
What the rule cost when it was applied (§3.6): `haq_index` left the scoring set (IHME is
`license_ok_commercial: false`, decided); four health indicators moved from WHO GHO to
World Bank mirrors; and `hale` was **the one scored indicator the rule did not clear**,
because it has no mirror. That exception was named on the §15 checklist rather than
tolerated silently, and **it was closed at the freeze**: REVIEW D10-B demoted `hale` to
display-only and moved its 0.14 to `life_expectancy` (§3.2, health). The health pillar now
carries no restricted source.
**The claim is now absolute, and it took two versions to earn.** At `0.1.1`, **zero scored
indicators sit on an uncleared source** — every id carrying weight is fed only by carriers
that are `license_ok_commercial: true` **and** `status: verified`, and the half-2 check
above asserts it on every commit.
That sentence could not be written before. `0.1.0` was minted with **three** exceptions;
D10-B closed `hale` in the same pass; the remaining **two** stood for the roughly one hour
`0.1.0` was the current version:
| indicator | pillar | weight at `0.1.0` | carrier | `license_ok_commercial` | rows staged |
|---|---|---|---|---|---|
| `tax_burden` | `legal_access` (Access) | 0.12 | `heritage_efi` | review | 6,951 |
| `freedom_house_score` | `stability` (FCS) | 0.04 | `freedom_house` | review | 7,095 |
`tax_burden` was already named on the §15.2 **[HUMAN]** list. `freedom_house_score` was
**not on any list** — `sources.yaml` states "until it returns, FIW feeds nothing scored"
(ISSUE-105) and it was in fact scored. Neither could be demoted inside `0.1.0`, because
demoting either re-weights a pillar and that is a **[HUMAN]** gate (§12, CLAUDE.md §3.10)
which D10 did not take. **REVIEW D12 took it**, and `0.1.1` demoted both (§16). ISSUE-132 is
closed.
**`0.1.0`'s record is not edited to match.** §15 still says the freeze shipped two
exposures, because it did, and a version's record is what was true when it was minted. What
changed is the version, not the history.
**What it cost the second time it was applied (§3.6.4, 2026-08-30).** The Phase-1 registry
was written against the source ids the scout *expected* to verify. When the scout actually
fetched the terms, eleven of them did not support the licence `03_DATA_SOURCES.md` claimed.
Rule 7 then removed **six scored indicators** — `terrorism_index`, `egov_index`,
`english_proficiency` (from both axes), `vdem_liberal_democracy`, `property_rights_index`
— and re-pointed five more onto World Bank mirrors. Two of the six are display-only in the
ordinary sense; two are *retired*, meaning not scored and not displayed either, because
their licences make display itself the prohibited act. **A rule that only ever costs
nothing is not a rule**, and this is the pass where it cost something: Access/fit is down to
two indicators, `legal_access` is 0.76 dependent on an unmerged PR, and the safety pillar
lost its terrorism signal.
8. **One carrier per indicator per country** — added at `0.2.0` (ballot A2). Where a
scored indicator's quantity is published by both a regional and a national carrier,
all units of a given country read the same carrier: the regional carrier wherever the
country is covered by it, the national carrier everywhere else — never both inside one
country. The two carriers may define the quantity differently, and a within-country
mix would publish the definition gap as if it were geography: OECD counts Mexico's
*practising* physicians (≈1.02/1k) where WDI relays the *licensed* count (≈2.59/1k),
a −60.7% wedge, and Germany's regional hospital-bed series undercounts the national
one by 23.5% on bed-type scope. Enforced by a staging validator: one `source_id` per
(indicator, country).
### 3.2 The indicator set (as amended at `0.2.0` — ballot A1/A2/B1/C1)
Weights are `weight_in_pillar`; each pillar's default-active weights sum to 1.0.
`dir` is `lower` when lower raw values are better.
**climate** (horizon-eligible)
| indicator_id | dir | w | note |
|---|---|---|---|
| `tasmax_warmest_month` | lower | 0.16 | heat block, §3.3 |
| `wetbulb_warmest_month_proxy` | lower | 0.10 | **proxy**, §3.3 |
| `heat_degree_months_30c` | lower | 0.08 | heat block, §3.3 |
| `coastal_flood_slr_exposure` | lower | 0.14 | LECZ, §3.4 |
| `drought_risk_aqueduct` | lower | 0.12 | §3.4; `drought_spei` is deferred |
| `riverine_flood_100yr_exposure` | lower | **0.12** | supersedes `riverine_flood_exposure`; Aqueduct Floods v2 100-yr inundation, baseline + 2050 (RCP4.5→ssp245, RCP8.5→ssp585, flagged); anchored |
| `riverine_flood_exposure` | lower | 0.0 | **deprecated** — `superseded_by: [riverine_flood_100yr_exposure]`; the annual-expected quantity has no projection |
| `wildfire_fwi` | lower | 0.10 | |
| `precip_variability` | lower | 0.06 | §3.4 |
| `cyclone_track_density` | lower | 0.06 | STORM (now) + STORM-C (2050, **ssp585 only**); ssp245 falls back to `now`, `no_projection`; anchored |
| `landslide_exposure` | lower | **display-only, 0.0** | NASA susceptibility, share of population in the two highest classes; projected variant (× rx5day change) context-only until a 0.3.x weights ballot |
| `local_warming_delta_c` | lower | **display-only, 0.0** | unit `tas` delta vs 1995–2014 beside the global delta from the same ensemble; copy rule §8 |
| `climate_analog_shift_km` | lower | 0.03 | |
| `disaster_loss_index` | lower | 0.03 | |
**water** (horizon-eligible)
| indicator_id | dir | w | note |
|---|---|---|---|
| `baseline_water_stress` | lower | 0.22 | |
| `projected_water_stress` | lower | 0.20 | Aqueduct 4.0 future water stress at 2035/2050 × both scenarios, **and the Aqueduct baseline at `now`** (`quality_flag: baseline_as_now`) so both horizons score the same indicator set — **ballot A5 retired C2.4's "no `now` value by construction"** (§4.4). Anchored, sharing `baseline_water_stress`'s curve. `bau` is Aqueduct's SSP3-7.0, read as `ssp245` and flagged `scenario_offset_ssp370` (ISSUE-148) |
| `groundwater_depletion_trend` | lower | 0.16 | |
| `renewable_freshwater_per_capita` | higher | 0.14 | |
| `safely_managed_drinking_water_pct` | higher | 0.14 | |
| `water_interannual_variability` | lower | 0.08 | |
| `desalination_capacity_per_capita` | higher | 0.06 | |
`desalination_capacity_per_capita` is the **compensator** REVIEW §1 Q2 requires: it
lives *inside* the Water pillar, so an arid coastal city with desalination never reaches
the veto threshold in the first place. Compensation is a pillar-level concept in this
methodology; the composite has no compensation mechanism at all.
**stability** (institutional only)
| indicator_id | dir | w |
|---|---|---|
| `wgi_political_stability` | higher | **0.30** |
| `wgi_rule_of_law` | higher | **0.30** |
| `wgi_control_of_corruption` | higher | **display-only, 0.0** — ρ 0.948 with rule_of_law; LOO rank ρ 0.99974 (ballot B1) |
| `wgi_government_effectiveness` | higher | **display-only, 0.0** — ρ 0.91 with the family; LOO rank ρ 0.99972 (ballot B1) |
| `freedom_house_score` | higher | **display-only, 0.0** — `review` carrier, §16 (D12) |
| `vdem_liberal_democracy` | higher | **display-only, 0.0** — CC-BY-SA, §3.6.4 |
| `debt_to_gdp` | lower | 0.10 |
| `inflation_volatility` | lower | 0.08 |
| `gdp_per_capita_growth` | higher | 0.08 |
| `reserves_months_imports` | higher | 0.06 |
| `fx_volatility_usd` | lower | 0.04 |
| `sovereign_rating` | higher | 0.04 |
| `fragile_states_index` | lower | **display-only, 0.0** |
The democratic-institutions family is now **four World Bank governance series and nothing
else** (§16). V-Dem is display-only pending counsel on share-alike; Freedom House is
display-only pending its permission reply. The publication-lag argument that earned FIW its
0.04 (§3.6.2) was never refuted — it simply cannot be paid for under rule 7 yet. At `0.2.0`
(ballot B1) the WGI family itself thinned to its two most distinct members: corruption and
government effectiveness are measured duplication (pairwise ρ 0.82–0.95 within the quartet,
`control_of_corruption` ~ `rule_of_law` at 0.948, LOO rank ρ ≥ 0.9997 per removal), so their
0.26 redistributed within the family and all four series stay on the card.
**health**
| indicator_id | dir | w | note |
|---|---|---|---|
| `life_expectancy_at_birth` | higher | **0.31** | supersedes `life_expectancy`; OECD TL2 where covered, national fallback (rule 8); gave 0.08 to `dependency_ratio` (ballot C1) |
| `dependency_ratio` | lower | **0.08** | UN WPP 2024, one carrier (rule 8); `now` = the 2023 estimate (WPP2024's last estimate year), 2050 = medium variant (`wpp_medium`); L0-native, inherited below with q 0.50; **"2050 demographic horizon (UN medium)"** |
| `old_age_dependency_ratio` | lower | 0.0 | **display-only** — WPP, the explanatory series |
| `pm25_surface_annual` | lower | **0.21** | supersedes `pm25_annual`; ACAG satellite surface, zonal per unit |
| `uhc_service_coverage_index` | higher | 0.19 | WDI mirror; took `haq_index`'s weight |
| `physicians_per_1000` | higher | 0.13 | supersedes `physicians_per_1k`; TL2 + national fallback (rule 8) |
| `hospital_beds_per_1000` | higher | 0.08 | supersedes `hospital_beds_per_1k`; TL2 + national fallback (rule 8) |
| `hale` | higher | 0.0 | **display-only** — WHO GHO licence, REVIEW D10-B |
| `haq_index` | higher | 0.0 | **display-only** — IHME licence, §3.6.1 |
| `pm25_annual` | lower | 0.0 | **deprecated** — `superseded_by: [pm25_surface_annual]`; quantity changed (modeled national exposure → measured surface concentration), see the 0.2.0 disclosure |
*(The four deprecated ids — `pm25_annual`, `life_expectancy`, `physicians_per_1k`,
`hospital_beds_per_1k`, plus `homicide_rate` below — move to `deprecated_indicators` in the
config with `superseded_by`, empty `sources` and empty `expected_sources`, per §3.6.6.)*
*`hale` was **demoted to display-only at the `0.1.0` freeze** by REVIEW D10-B (Eric,
2026-08-30), for licence reasons and no other: WHO GHO is CC-BY-NC-SA 3.0 IGO
(`license_ok_commercial: review`) and HALE has no World Bank mirror, so §3.1 rule 7 could
not clear it. Its 0.14 went to `life_expectancy`, which asks the same longevity question of
a CC-BY-3.0-IGO carrier; the two correlate above 0.98 across countries, so the family keeps
its shape while carrying it on one leg instead of two. **HALE is still staged for 7,097 units
and still shown on the card, marked** — display is what the licence permits. It is the first
item queued for restoration if WHO commercial permission arrives, which would be a major
bump and a **[HUMAN]** gate (§12), not an automatic reversal.*
*`blue_zone_proximity` (0.04) was **removed** by Eric at the manual-table review
(2026-08-29/30): the Blue Zones research is contested (record-keeping and pension-fraud
critiques), and the product's own thesis supersedes the concept — measured life-outcome
data over folklore geography. It was never staged, so no published score changes; its
0.04 went +0.01 each to the four cleanly-licensed carriers above, deliberately not to
`hale`. This is a decided item, not a §15 open one.*
**infrastructure**
| indicator_id | dir | w |
|---|---|---|
| `electricity_access_pct` | higher | 0.17 |
| `grid_reliability` | higher | 0.15 |
| `logistics_performance_index` | higher | 0.15 |
| `renewable_energy_share` | higher | 0.13 |
| `internet_users_pct` | higher | 0.10 |
| `airport_hub_distance_km` | lower | 0.09 |
| `broadband_speed_fixed` | higher | 0.06 |
| `road_density` | higher | 0.06 |
| `mobile_subscriptions_per_100` | higher | 0.05 |
| `mobile_broadband_speed` | higher | 0.04 |
| `egov_index` | higher | **retired, 0.0** — UN terms, §3.6.4 |
The connectivity block (`internet_users_pct`, `broadband_speed_fixed`,
`mobile_subscriptions_per_100`, `mobile_broadband_speed`) is §3.5's four-way subdivision of
what was one 0.24 speed pairing; it totals 0.25 after absorbing its share of `egov_index`.
**safety**
| indicator_id | dir | w | note |
|---|---|---|---|
| `homicides_per_100k` | lower | **0.51** | supersedes `homicide_rate`; TL2 + national fallback (rule 8); took its share of seismic's 0.20 (ballot C1) |
| `conflict_events_per_100k` | lower | **0.39** | always on; took its share of seismic's 0.20 |
| `volcanic_hazard` | lower | **0.10** | always on (natural hazard); took its share of seismic's 0.20 |
| `seismic_hazard` | lower | 0.0 | **display-only** — 244/7,370 units, centroid sample at the wrong return period (ISSUE-110, ISSUE-123); restores on a cleared, spec-conformant global grid — major bump, [HUMAN] |
| `distance_to_active_conflict_km` | higher | 0.12 | **gated:** `resilience_first` |
| `firearm_legality` | higher | 0.10 | **gated:** `resilience_first` |
| `terrorism_index` | lower | 0.0 | **retired** — IEP licence, §3.6.4 |
`global_peace_index` was **removed from the model** on 2026-08-30 (§3.1 rule 3 table).
**community**
| indicator_id | dir | w |
|---|---|---|
| `happiness_cantril` | higher | 0.25 |
| `social_support` | higher | 0.18 |
| `mean_years_schooling` | higher | 0.16 |
| `protected_area_access` | higher | 0.16 |
| `human_capital_index` | higher | 0.14 |
| `expat_community_density` | higher | 0.11 |
| `english_proficiency` | higher | **display-only, 0.0** — no licence, §3.6.4 |
| `oecd_regional_wellbeing_index` | higher | **display-only, 0.0** |
| `tertiary_attainment_25_64_pct` | higher | **display-only, 0.0** — OECD TL2, new quantity; weighting deferred to its own ballot |
| `voter_turnout_pct` | higher | **display-only, 0.0** — OECD TL2, new quantity; weighting deferred to its own ballot |
**food**
Two blocks since Phase 1 (§3.6.3): **land capacity 0.70** and **supply security 0.30**.
| indicator_id | dir | w | block |
|---|---|---|---|
| `crop_yield_change_2050_pct` | higher | **0.17** | capacity — ISIMIP3b GGCMI, production-weighted maize/wheat/rice/soy, ensemble median ≥ 5 crop models × ≥ 3 GCMs; ssp585 direct, **ssp370 → ssp245 flagged `scenario_offset_ssp370`**; "2050 climate horizon" |
| `agro_suitability_gaez` | higher | **0.0** | capacity — **display-only**, barred (GAEZ CC-BY-NC-SA); weight passed to `crop_yield_change_2050_pct` (ballot C2) |
| `growing_season_length` | higher | 0.15 | capacity |
| `arable_land_per_capita` | higher | 0.14 | capacity |
| `soil_quality` | higher | 0.12 | capacity |
| `tree_cover_loss_trend` | lower | 0.12 | capacity |
| `food_import_dependence` | lower | 0.16 | supply security |
| `prevalence_of_undernourishment` | lower | 0.14 | supply security |
| `cereal_yield_trend_5y` | higher | **display-only, 0.0** | capacity, trend |
*(Global Food Security Index is absent by rule 3.1.3 — composite of our own inputs — and
REVIEW D3 confirms that is a double-count ruling, not a licence one.)*
**Access — affordability**
| indicator_id | dir | w |
|---|---|---|
| `price_to_income_ratio` | lower | 0.28 |
| `cost_of_living_index` | lower | 0.24 |
| `real_house_price_trend_5y` | lower | 0.14 |
| `construction_cost_index` | lower | 0.14 |
| `insurance_availability` | higher | 0.12 |
| `rent_yield` | higher | 0.08 |
**Access — legal_access**
| indicator_id | dir | w |
|---|---|---|
| `foreign_ownership_rules` | higher | **0.46** |
| `visa_pathways` | higher | **0.40** |
| `passport_mobility_henley` | higher | **0.14** |
| `tax_burden` | lower | **display-only, 0.0** — `review` carrier, §16 (D12) |
| `property_rights_index` | higher | **display-only, 0.0** — IPRI/Heritage licences, §3.6.4 |
**`legal_access` scores nothing at `0.1.1`, and that is a licence outcome, not a modelling
one.** 0.86 of the pillar sits on the two manual tables of §3.7, which are approved but not
staged into `indicator_values`; `passport_mobility_henley` holds the remaining 0.14 with no
rows. `tax_burden` was the only member with staged values — 6,951 units — and D12 demoted it
(§16). So the pillar's score is null for every unit, its confidence reads 0, and the Access
composite renormalizes over `affordability` and `fit` alone (§4.3, §9). **Read the Access axis
at `0.1.1` as two-thirds of a design**, and read §16 for what it costs.
**Access — fit**
| indicator_id | dir | w |
|---|---|---|
| `winter_severity` | lower | 0.55 |
| `timezone_overlap_us_hours` | higher | 0.45 |
| `english_proficiency` | higher | **display-only, 0.0** — no licence, §3.6.4 |
**Fit is now two in-house derivations** — a temperature transform and a time-zone overlap —
carrying a 20-weight Access pillar between them. That is thin enough to be a [HUMAN] item
(§15.2), not a rounding detail.
Dominant religion/culture flags from brief §3.2 are **informational only** and are not
indicators. They do not enter any score.
### 3.3 The heat block, and two deferred indicators
The brief specified heat as **threshold-day counts** — `heat_days_35c`
(`count(tasmax ≥ 308.15 K)` per year) and `wetbulb_days_31c`. The Scout established that
neither is buildable at v0.1 (`docs/DATA_ISSUES.md` **ISSUE-005**): the NEX-GDDP-CMIP6
COG product that actually exists is an **ensemble-median monthly mean**, and a count of
days above a threshold is not recoverable from a monthly mean. Daily COGs exist for two
models only — below the ≥ 10-model ensemble rule in architecture §3.4.
Rather than ship a threshold-day number we cannot compute, or a 2-model ensemble we
would have to disclaim, v0.1 **redefines the heat block as quantities the monthly
ensemble medians genuinely support**:
| indicator | what it is | source form |
|---|---|---|
| `tasmax_warmest_month` | mean daily maximum temperature of the warmest month, °C | monthly ensemble-median `tasmax`, max over months |
| `heat_degree_months_30c` | Σ over months of `max(0, tasmax_month − 30 °C)`, °C·months | monthly ensemble-median `tasmax` |
| `wetbulb_warmest_month_proxy` | wet-bulb temperature of the warmest month, computed from **monthly mean** `tasmax` + `hurs` | monthly ensemble-median `tasmax`, `hurs` |
**The limitation, stated plainly** (and it must be stated on `/methodology` too):
- A monthly mean **understates extremes**. Two places with the same warmest-month mean
can have very different numbers of genuinely dangerous days; the heat block measures
sustained seasonal heat, not peak lethality.
- `wetbulb_warmest_month_proxy` is the weaker of the three, because humid heat is an
extreme-value phenomenon and a monthly mean of it is a systematic underestimate. It is
therefore flagged `proxy: true`, **down-weighted to 0.10** relative to the 0.15 the
brief gave the threshold-day version, and must carry a proxy marker in the UI wherever
it is shown.
- The heat block totals 0.34 of the Climate pillar (the brief's threshold-day pair
totalled 0.30), split across three partly-redundant monthly quantities rather than
concentrated in two indicators we would be inventing.
**Deferred, not renamed.** Indicator ids are stable forever (CLAUDE.md §4), so
`heat_days_35c` and `wetbulb_days_31c` are **reserved** in
`deferred_indicators` — they may never be reused for a different quantity, and they come
back under their own ids in Phase 1 when the daily NetCDF archive is ingested on the
rented raster box. Restoring them is an **indicator-set change ⇒ major bump** (§12).
`drought_spei` joins them for the same reason at §3.4.3.
### 3.4 Three climate transforms, written down (2026-08-29)
The [geo] agent built the Phase-0 climate and water layers and handed four questions back
to this role (`docs/plans/2026-08-28-geo.md` §6, ISSUE-G-001/002/004/006), because three
of these indicator ids were frozen here **with no transform attached** — a bare id is an
invitation to build a different quantity than the one the model means. The transforms are
therefore normative from here on. All four decisions are recorded in `docs/CHANGELOG.md`
under `0.1.0-draft` and resolved in `docs/DATA_ISSUES.md`.
#### 3.4.1 `coastal_flood_slr_exposure` — low-elevation coastal zone (LECZ)
> The share of the unit's population living in reference-grid cells that are **both**
> below (10 m + the horizon's sea-level increment) mean surface elevation **and** within
> 50 km of the coastline.
Population weights are GHSL `R2023A-E2025` (1 km, sum-resampled to the 0.1° reference
grid); elevation is ETOPO 2022 60″ surface; the coastline is Natural Earth 10 m. Range
`[0, 1]`, `lower_is_better`. This is the standard LECZ construction and it is what the id
says: population *in the zone* is exposure.
It is **not** Aqueduct's `cfr_raw`, which was the obvious source and which fails on data
rather than on method: the Aqueduct sub-basin containing Miami carries `cfr_raw = 0.0` /
`cfr_cat = −1` (no data), which ranked Miami's coastal-flood exposure below Seattle's.
`cfr_raw` returns if a future Aqueduct release fixes that basin.
##### The horizon increment (added 2026-08-29)
Until this amendment the indicator was `now`-only, which meant **the sea did not rise
between horizons** — the one indicator carrying the sea-level story was frozen at its
present value while the model claimed to project to 2050 (`docs/QA_REPORT.md` 2026-08-29,
anchor 5). The threshold now rises with the horizon by the IPCC AR6 medium-confidence
**global-mean** sea-level median, relative to the 1995–2014 baseline:
| horizon | scenario | increment | threshold | provenance |
|---|---|---|---|---|
| `now` | `obs` | 0.00 m | 10.00 m | — |
| `2035` | `ssp245` | 0.12 m | 10.12 m | interpolated |
| `2035` | `ssp585` | 0.13 m | 10.13 m | interpolated |
| `2050` | `ssp245` | **0.20 m** | 10.20 m | AR6 Table 9.9 median |
| `2050` | `ssp585` | **0.23 m** | 10.23 m | AR6 Table 9.9 median |
Source: **IPCC AR6 WG1 Chapter 9, Table 9.9**, "Total" rows, unshaded (medium-confidence)
columns — 2030: SSP2-4.5 0.09 (0.08–0.12) m, SSP5-8.5 0.10 (0.09–0.12) m; 2050: SSP2-4.5
0.20 (0.17–0.26) m, SSP5-8.5 0.23 (0.20–0.29) m. Registered as source `ipcc_ar6_slr`; the
verbatim extract and its manifest are in `data/raw/ipcc_ar6_slr/wg1-ch9-table-9.9/`.
Table 9.9 publishes no 2040 row, so **2035 is linearly interpolated** between the 2030 and
2050 medians. GMSL is accelerating, so linear interpolation over that span slightly
*over*-states 2035, by roughly a centimetre — four orders of magnitude below what a 0.1°
grid of 0.1°-mean elevations can discriminate.
**Limitation — global mean, not regional.** Every coast gets the same increment. AR6's
*regional* relative sea level (vertical land motion, gravitational fingerprints, ocean
dynamics) departs from the global mean by tens of centimetres in places, and in opposite
directions: subsiding deltas get more, post-glacial rebound coasts get less. Upgrading to
regional projections via the NASA AR6 sea-level tool named in `docs/03_DATA_SOURCES.md` is
**Phase 1**. Until then this indicator ranks coasts by *how much population sits near sea
level*, adjusted by a uniform global shift — not by how much the sea will rise at each one.
**Limitation — the signal is small by construction, and was measured rather than
asserted.** A 0.20 m rise against a 10 m threshold is a 2 % move that only reclassifies
cells whose 0.1°-mean elevation falls in [10.00, 10.20] m. Measured over the Phase-0 320:
**43 units move at all** between `now` and 2050/SSP2-4.5, by a mean of **+0.0047** and a
maximum of **+0.057**. Units already saturated cannot move — **Miami sits at 1.000 at every
horizon** — and under percentile-rank normalization against a frozen reference (§4.1) a
unit already ranked worst in its cohort *cannot be made worse by raising the sea*. So this
amendment gives the indicator a real horizon signal, but it does not and cannot make
sea-level rise the dominant term at v0.1.
**Limitation — connectivity.** At 0.1° (≈ 11 km) the LECZ's usual hydrological-connectivity
test degrades to a straight-line distance test, so a low cell behind a ridge or a levee
still counts. The indicator measures *where people live relative to sea level and the
coast*, not defended flood risk.
#### 3.4.2 `precip_variability` — CV of the monthly totals in the window
> The coefficient of variation (σ/μ) of the monthly precipitation totals across the
> climate window — 60 monthly values at the Phase-0 5-year window.
**Limitation, stated because it changes how the number reads.** This conflates the
*seasonal cycle* with *interannual* variability: a Mediterranean climate, whose rain is
reliable but concentrated, scores "variable". Blessed for v0.1 anyway, because the
interannual-only alternative (CV of the 5 annual totals) is an estimator over `n = 5` and
would be noise, and because seasonal concentration is itself a storage-and-supply signal.
At v0.2, once the 20-year window exists (§3.4.4), it splits into `precip_seasonality_cv`
and `precip_interannual_cv` and `precip_variability` is **deprecated, not redefined**.
Left undefined it would be a value nobody can check; the weight is 0.06 either way.
#### 3.4.3 `drought_spei` is deferred; the Phase-0 quantity is `drought_risk_aqueduct`
The only drought source available at Phase 0 is Aqueduct 4.0's `drr_raw`, which is
Aqueduct's own **hazard × exposure × vulnerability composite** — not a Standardized
Precipitation-Evapotranspiration Index. Publishing it under `drought_spei` would put a
different quantity behind a named index. So:
- `drought_spei` → `deferred_indicators`, reserved for a real SPEI, `superseded_by:
[drought_risk_aqueduct]`, revisit in Phase 1 from a 20-year precipitation + PET window.
- **`drought_risk_aqueduct`** — climate, `lower_is_better`, `weight_in_pillar` 0.12
(the weight `drought_spei` held, so the pillar still sums to 1.0). Population-weighted
`drr_raw`, now/obs only.
**Declared caveat.** Aqueduct's DRR carries a socio-economic vulnerability layer, so it is
not a pure physical hazard and sits closer to rule §3.1.3 (composites of our own inputs)
than the rest of the climate block. It is scored at v0.1 because the alternative is no
drought signal at all; whether it should be display-only is a v0.2 review item.
#### 3.4.4 Phase-0 climate windows are **5 years**, not the 20 the architecture specifies
Architecture §3.4 specifies a 20-year window. Phase 0 uses five, for download and
wall-clock economy on a slice whose acceptance test is a *relative* ranking:
| horizon | scenario | window |
|---|---|---|
| `now` | `obs` | 2010–2014 (`historical`) |
| `2035` | `ssp245`, `ssp585` | 2033–2037 |
| `2050` | `ssp245`, `ssp585` | 2048–2052 |
**Consequence, and it is not cosmetic.** A 5-year mean of an ensemble median is a noisy
estimate of a climate normal. Cross-unit *rankings* — which is all §4 consumes — are
robust to that noise because the noise is largely common-mode across units in a region;
**absolute** values are not, and must not be quoted as climate normals. Anything derived
from a variance is the most affected, which is `precip_variability` (§3.4.2). Widened to
20 years on the rented raster box (REVIEW §6.5) before the `0.1.0` freeze.
Aqueduct's future scenarios are also mapped rather than native: `bau` → `ssp245`,
`pes` → `ssp585`, and Aqueduct's 2030 horizon is carried as `2035` with
`quality_flag = horizon_offset_2030` on every affected row.
---
### 3.5 Infrastructure, safety, food and affordability get their first indicators (2026-08-29)
The first real-data run scored **three FCS pillars for no unit at all** — `infrastructure`
(10 weight points), `safety` (6) and `food` (5) — and left the whole Access axis null, which
made anchor 6b unevaluable (`docs/QA_REPORT.md` 2026-08-29). Six World Bank WDI series close
the worst of that gap. All are `now` + `trend` quantities and are **never projected**
(CLAUDE.md §3 rule 5); all carry `as_of` at the true observation year so the recency factor
(§7) discounts a stale value rather than the pipeline hiding it.
| WDI series | indicator id | pillar | direction | weight in pillar |
|---|---|---|---|---|
| `EG.ELC.ACCS.ZS` | `electricity_access_pct` | infrastructure | higher | 0.16 *(existing)* |
| `IT.NET.USER.ZS` | `internet_users_pct` | infrastructure | higher | **0.09 (new)** |
| `IT.CEL.SETS.P2` | `mobile_subscriptions_per_100` | infrastructure | higher | **0.05 (new)** |
| `VC.IHR.PSRC.P5` | `homicide_rate` | safety | lower | 0.32 *(existing)* |
| `AG.LND.ARBL.HA.PC` | `arable_land_per_capita` | food | higher | 0.16 *(existing)* |
| `PA.NUS.GDP.PLI` | `price_level_index_gdp` | access/affordability | lower | **0.12 (new)** |
#### 3.5.1 Where the weight for the three new ids came from
Pillar weights must sum to exactly 1.0, so three new ids need weight from somewhere. In both
cases it comes **from inside the same family**, so no pillar's shape changes and no existing
indicator outside the family moves at all.
**Infrastructure — the connectivity block keeps its 0.24 and is subdivided four ways:**
| id | before | after |
|---|---|---|
| `internet_users_pct` | — | **0.09** |
| `broadband_speed_fixed` | 0.14 | 0.06 |
| `mobile_subscriptions_per_100` | — | **0.05** |
| `mobile_broadband_speed` | 0.10 | 0.04 |
| *(block total)* | **0.24** | **0.24** |
Rationale: **penetration outranks speed as a resilience measure.** Whether a population is
connected at all is a structural fact about a place; headline speed is a quality refinement
on top of it, and is an Ookla-class quantity Phase 0 cannot source at all. Penetration takes
0.14 of the block's 0.24; speed keeps 0.10.
**Access/affordability — the cost family keeps its 0.24, split evenly:**
`cost_of_living_index` 0.24 → 0.12, `price_level_index_gdp` 0.12. The consumer-basket index
stays in the model, unsourced, because its natural source (Numbeo) is barred by CLAUDE.md
§3.4 and Phase 1 may licence one. The national price level is the same question — *how
expensive is this place* — asked at the economy level rather than the basket level.
**Safety and food take no weight change**: both new series land on ids the model already
carried at their existing weights.
#### 3.5.2 Two series that did not survive verification, and why that matters
**`PA.NUS.PPPC.RF` is archived.** The World Bank API answers *"The indicator was not found.
It may have been deleted or archived."* The live member of the same family is
`PA.NUS.GDP.PLI` — price level index (GDP), US = 100 — which is the same quantity (PPP
conversion factor ÷ market exchange rate) expressed as an index rather than a ratio. The id
registered is `price_level_index_gdp`, named for what it actually is. (ISSUE-015.)
**`TM.VAL.FOOD.ZS.UN` is rejected for `food_import_dependence`.** It is food as a share of
*merchandise imports*, so the denominator is the whole import basket rather than food
consumption. It ranks New Zealand (12.8 %) as more food-dependent than the United States
(6.9 %), and Switzerland (4.6 %) as barely dependent while it imports roughly half its
calories. Publishing it under `food_import_dependence` would put a different quantity behind
a named id — precisely the error §3.4.3 refused for `drought_spei`. **Food therefore ships
with `arable_land_per_capita` alone**, and the gap is logged (ISSUE-016).
#### 3.5.3 What the food pillar does and does not measure
Worth stating explicitly, because the Phase-0 result is counterintuitive and is *not* a bug.
Every indicator in the food pillar — `agro_suitability_gaez`, `growing_season_length`,
`arable_land_per_capita`, `soil_quality`, `tree_cover_loss_trend` — measures **the land's
capacity to produce food**, not whether people are currently fed. Sudan consequently scores
**86.4** on food (0.420 ha arable land per person, second only to the USA) while in famine,
and Egypt scores **4.5** (0.027 ha/person, worst of the eleven).
That is the intended division of labour: a famine driven by conflict and distribution
failure is routed through `stability` (Sudan 15.97) and `health` (17.13), where Sudan is duly
vetoed, not through `food`. A food-security *outcome* indicator (undernourishment prevalence,
IPC phase) is a different quantity from land capacity and would need its own id and its own
weight — a Phase-1 addition, not a Phase-0 patch. It was deliberately **not** added to close
anchor 2, which would have been tuning to force an intuitive answer.
#### 3.5.4 Coverage holes recorded in advance
`VC.IHR.PSRC.P5` has no observation after **2008 for Sudan** and 2013 for Yemen;
`IT.NET.USER.ZS` stops at 2017 for Sudan and 2019 for Yemen. The fetch window is therefore
`1990:2025` for this block so the value exists at all. **Sudan's 2008 homicide rate predates
its civil war entirely**, so the safety pillar *understates* Sudan — recorded here so the
anchor-2 result is read honestly rather than credited to this amendment.
---
### 3.6 The Phase-1 registry expansion (2026-08-29)
Phase 0 shipped a model in which **nine of the sixty-two scored indicators had a value**. This
section registers the Phase-1 set — the ids, weights and dispositions the [pipeline] agent
stages against in wave B. Everything here is a **draft amendment** (the version is unfrozen);
each would be a **major** change after the freeze, so each is written down for the freeze diff.
**The weight discipline, stated once.** Where an id is added, its weight comes from **inside its
own family**, so the family total is unchanged and no indicator outside the family moves at all.
There is exactly one exception — the food pillar in §3.6.3 — and it is called a reshape rather
than disguised as a family split.
#### 3.6.1 Health: four WHO series move to World Bank mirrors, and one leaves the model
Rule §3.1.7 bit hardest here. **0.54 of the health pillar — PM2.5, HALE, physicians and hospital
beds — was fed by WHO GHO**, which is CC-BY-NC-SA 3.0 IGO and therefore `review`. A further
0.18 sat on `haq_index`, whose source (IHME) is `license_ok_commercial: false` and was never
staged at all. REVIEW §3 row 8 named the remedy in advance: *"World Bank WDI mirrors LE,
physicians, beds (CC-BY) — safe harbor for Pro tier."*
| id | was | now | weight |
|---|---|---|---|
| `physicians_per_1k` | GHO `HWF_0001` ÷ 10 | WDI **`SH.MED.PHYS.ZS`** | 0.12 (unchanged) |
| `hospital_beds_per_1k` | GHO `WHS6_102` ÷ 10 | WDI **`SH.MED.BEDS.ZS`** | 0.08 (unchanged) |
| `pm25_annual` | GHO `SDGPM25` | WDI **`EN.ATM.PM25.MC.M3`** | 0.20 (unchanged) |
| `uhc_service_coverage_index` | GHO `UHC_INDEX_REPORTED`, display-only | WDI **`SH.UHC.SRVS.CV.XD`**, **scored** | **0.0 → 0.18** |
| `haq_index` | scored, never sourced | **display-only** | **0.18 → 0.0** |
| `hale` | GHO `WHOSIS_000002` | unchanged — **no WDI mirror exists** | 0.14 (unchanged) |
**Same ids, no deprecation, and that is the rule rather than convenience.** An id is deprecated
when the *quantity* behind it changes (`drought_spei` §3.4.3, `food_import_dependence`
ISSUE-016). A **carrier** change is not a redefinition: physicians per 1,000 is physicians per
1,000 whether WHO or the World Bank republishes the national return. The numbers will move
slightly — the two publishers carry different vintages, and for PM2.5 different models (WHO's
own surface vs. the World Bank's republication of the Global Burden of Disease estimate) — and
that difference is recorded in each spec rather than smoothed over.
**The health-system family total is unchanged at 0.38.** `haq_index` held 0.18 and could never
use it; that 0.18 goes to `uhc_service_coverage_index`, and physicians (0.12) and beds (0.08) do
not move. **`uhc` is the weaker measure and the better-licensed one**, and the trade is explicit:
IHME's HAQ index is built from *amenable mortality* — deaths that should not have happened given
timely care — while the UHC index is built from service-coverage *inputs*. We are substituting a
coverage measure for an outcome measure because the outcome measure cannot be published.
**The staged rows have not caught up.** These are spec changes; the Phase-0 rows still carry
`source_id: who_gho`. `uv run marts score` now reports that gap by name as
`staged source_id != the spec's sources`, and it must read **0** before a version is published
(§15).
#### 3.6.2 Stability, water, safety, infrastructure, community
**Read this table with §3.6.4.** It records what was *planned* on 2026-08-29 against the source
ids the scout expected to verify. Nine of its rows were overtaken the next day when the terms
were actually read; each is struck through and re-stated in §3.6.4 rather than quietly edited,
because the freeze diff needs the difference between what was planned and what shipped.
| pillar | change | weight |
|---|---|---|
| stability | **`freedom_house_score` added** | 0.04, from `vdem_liberal_democracy` 0.12 → 0.08; **democratic-institutions family total unchanged at 0.12** — *superseded: V-Dem is now display-only, §3.6.4* |
| stability | `fx_volatility_usd` sourced (Frankfurter/ECB reference rates) | 0.04, unchanged |
| water | `renewable_freshwater_per_capita`, `safely_managed_drinking_water_pct` sourced — *AQUASTAT and JMP replaced by the `wdi_water` mirrors, §3.6.4* | 0.14 + 0.14, unchanged |
| safety | `conflict_event_density_50km` sourced (**UCDP GED**, not ACLED — ACLED is licence-blocked) | 0.24, unchanged |
| safety | ~~`homicide_rate` primary carrier **WDI → UNODC**~~ — *cancelled: UNODC is non-commercial and no-derivatives; WDI is the permanent carrier, §3.6.4* | 0.32 → **0.41**, §3.6.4 |
| safety | `seismic_hazard` (USGS, **US-only**), `volcanic_hazard` (GVP, licence cleared 2026-08-30) sourced; both **always on** (§3.1 rule 2) | 0.20 + 0.08, unchanged |
| safety | ~~`terrorism_index` stays scored, `global_peace_index` stays display-only~~ — *both removed by the IEP licence, §3.1 rule 3 table and §3.6.4* | 0.16 → **0.0** / removed |
| infrastructure | `renewable_energy_share` (Ember), `logistics_performance_index` (WB LPI), `airport_hub_distance_km` (OurAirports) sourced; ~~`egov_index` (UN EGDI)~~ *retired, §3.6.4* | egov 0.06 → **0.0**, rest re-scaled |
| community | `social_support` sourced from **OECD Regional Well-Being's Community topic** — closes ISSUE-010 | 0.16, unchanged |
| community | `human_capital_index` (WB HCI) sourced; `mean_years_schooling` *re-pointed to the World Bank's UIS republication*; ~~`english_proficiency` (EF EPI)~~ *demoted, §3.6.4* | english 0.12 → **0.0**, rest re-scaled |
| community | `oecd_regional_wellbeing_index` added **display-only** (§3.1 rule 3) | 0.0 |
| access | ~~BIS RPPI (`real_house_price_trend_5y`)~~ — *superseded 2026-08-30: re-pointed to OECD `DSD_AN_HOUSE_PRICES` `MEASURE=RHP`, §3.6.5*; OECD price-to-income (**licence cleared 2026-08-30**), ICP upgrade path for `price_level_index_gdp`; ~~IPRI + Heritage `property_rights_index`~~ *display-only, §3.6.4*; `tax_burden` still Heritage-dependent | property rights 0.20 → **0.0** |
| access/fit | **`winter_severity` and `timezone_overlap_us_hours` become buildable in-house** from already-verified sources (NEX-GDDP monthly `tasmax`; Natural Earth time zones), so `fit` stops being an empty pillar — and, after `english_proficiency` was demoted, becomes the whole of it | 0.36 → **0.55**, 0.30 → **0.45** |
Two of these deserve their sentence:
- **Freedom House is not redundant with V-Dem** under rule 3, because rule 3 demotes an index
built from quantities *we* already score, and Freedom House is built from its own analyst
survey. It earns its 0.04 on **publication lag**: Freedom in the World lands months before the
V-Dem release, so a democratic reversal reaches the model a cycle earlier — which is exactly
the political-era drift REVIEW D6 wants visible on an annual cadence.
- **`social_support` is severely coverage-limited and is still the right source.** OECD Regional
Well-Being covers OECD members and a few partners, so roughly forty countries get a value and
the rest take the coverage penalty. The alternative — the raw Gallup social-support variable —
is not free, and the World Happiness Report publishes only the ladder *decomposition*
(ISSUE-010), which is a contribution in ladder points, not the variable.
#### 3.6.3 Food is reshaped, and the reshape is declared
Every indicator in the food pillar measured **the land's capacity to produce food** and none
measured whether anyone is fed — which is why Sudan scored 86.4 on food while in famine
(§3.5.3). §3.5.3 named the fix as a Phase-1 addition needing *"its own id and its own weight"*.
That is what this is.
| block | before | after |
|---|---|---|
| land capacity (agro-suitability, growing season, arable land, soil, tree-cover loss) | 0.82 | **0.70**, proportions preserved — every member rescaled by the same factor, so none moved relative to the others |
| supply security | 0.18 (`food_import_dependence` alone) | **0.30** — `food_import_dependence` 0.16 + **`prevalence_of_undernourishment` 0.14** |
`food_import_dependence` is also finally sourced correctly: FAOSTAT's **cereal import dependency
ratio**, whose denominator is domestic supply — the quantity the id claims — rather than WDI's
`TM.VAL.FOOD.ZS.UN`, which is food as a share of *merchandise imports* and was rejected as a
different quantity behind a named id (ISSUE-016).
**Read the motive correctly.** `prevalence_of_undernourishment` is defined by its quantity, not
by the anchor it happens to move. Anchor 2 (Yemen and Sudan at the bottom of the Phase-0 ten) is
a **strict xfail**, so if this addition flips it the test suite *reports* the flip — the flip is
not the point, and it is not evidence the addition was tuning. A 70/30 capacity/supply split is
the smallest change that gives the pillar a supply-security block at all; it is a **minor** bump
and a **[HUMAN]** item at the freeze (§15).
#### 3.6.4 The licence reconciliation (2026-08-30) — what the registry actually gets
§3.6.1–3.6.3 were written on 2026-08-29 against the source ids the scout said it would
register. The scout then fetched the terms and **eleven Phase-1 rows in
`docs/03_DATA_SOURCES.md` carry a licence the actual terms do not support** (DATA_ISSUES
ISSUE-101…115). This section reconciles the registry with those verdicts. Nothing here is a
methodological improvement; every line is a licence consequence, and the model is worse for
several of them. Recording that plainly is the point — the alternative is a score whose inputs
nobody may lawfully publish.
**Two sources were cleared** (browser-verified 2026-08-30, `sources.yaml` flipped to
`license_ok_commercial: true`):
| source | verdict | what it unblocks |
|---|---|---|
| `oecd_data_explorer` | OECD T&C §3 Data → Permitted Use: Data may be extracted, adapted, distributed and embedded **"for any purpose, even for commercial use"**, with a fixed attribution string and a pass-down acknowledgment obligation in any sublicence | `social_support`, `price_to_income_ratio`, `oecd_regional_wellbeing_index`, TL2 homicide |
| `gvp_volcanoes` | GVP Terms of Use: the database compilation is **a product of US-government employees**, so it is not eligible for copyright. Database only — not photographs, not the bulletins | `volcanic_hazard`, scored at 0.08 |
**Five indicators were re-pointed** — same id, same quantity, cleaner carrier, exactly the
WHO→World Bank manoeuvre of §3.6.1:
| id | was | now | why |
|---|---|---|---|
| `homicide_rate` | WDI, with UNODC planned as the Phase-1 primary | **WDI `VC.IHR.PSRC.P5`, permanently** | UNODC's terms are non-commercial **and** no-derivatives (ISSUE-109). WDI republishes the same UNODC series under CC-BY-4.0; we lose the newer vintage and the sex/age disaggregation |
| `renewable_freshwater_per_capita` | FAO AQUASTAT | **`wdi_water` `ER.H2O.INTR.PC`** | AQUASTAT has no bulk endpoint and unconfirmed terms (ISSUE-114); the WDI mirror is the same series, CC-BY-4.0 |
| `safely_managed_drinking_water_pct` | WHO/UNICEF JMP | **`wdi_water` `SH.H2O.SMDW.ZS`** | washdata.org publishes no licence instrument at all (ISSUE-114) |
| `mean_years_schooling` | UNESCO UIS API | **World Bank `UIS.EA.MEAN.1T6.AG25T99`** | UIS is CC-BY-SA-4.0 and share-alike is unresolved (ISSUE-104). The World Bank republishes the same series under CC-BY-4.0 — verified live on 2026-08-30, 185 economies, but **frozen at 2024-06-25** with most observations from 2016–2018, so it must be labelled with its vintage and carries no trend |
| `conflict_event_density_50km`, `seismic_hazard`, `volcanic_hazard` | pending | **`ucdp`, `usgs_seismic`, `gvp_volcanoes` registered** | all three are now verified and commercially cleared |
**Six indicators lost their weight.** Two are *retired* — a stronger disposition than
display-only, and one this methodology had not needed before: **an indicator whose licence
makes display itself the prohibited act cannot be shown as corroboration either.** Retired ids
keep their id, carry weight 0, and hold empty `sources` **and** empty `expected_sources`, which
is the enforceable half of the statement: no ingester may fetch them, so no value can be
staged, so nothing can be rendered.
| id | pillar · old w | disposition | the sentence that must travel with it |
|---|---|---|---|
| `terrorism_index` | safety · 0.16 | **retired** | *The Global Terrorism Index is sold under a commercial licence, and its publisher counts any use by a company — including showing it — as commercial use, so we do not carry it at all. Organised violence is measured directly from UCDP event data instead.* |
| `egov_index` | infrastructure · 0.06 | **retired** | *The UN's e-government survey is published under terms that permit neither commercial use nor derivative works, so we cannot score it or show it. Digital service delivery is a gap in this pillar, not a solved problem.* |
| `english_proficiency` | fit · 0.34 **and** community · 0.12 | display-only pending a licensed table | *EF publishes its English proficiency index as a report, with no data file and no licence, so we show it only where we can cite it directly.* |
| `vdem_liberal_democracy` | stability · 0.08 | display-only pending counsel | *V-Dem is share-alike licensed; until our lawyers confirm that a percentile rank is not a derivative work, we show it beside the score rather than inside it.* |
| `property_rights_index` | legal_access · 0.20 | display-only pending Heritage | *The two published property-rights indices are either barred from commercial reproduction or published with no licence at all, so we show a property-rights reading beside the Access score rather than inside it.* |
| `global_peace_index` | safety · 0.0 | **removed** | (was already display-only; see §3.1 rule 3) |
**Weight redistribution, all within family, all proportional to existing weights.** Every
pillar's default-active weights still sum to exactly 1.0 and no indicator outside an affected
family moved:
| pillar | redistribution |
|---|---|
| safety | GTI's 0.16 → `homicide_rate` 0.32 → **0.41**, `conflict_event_density_50km` 0.24 → **0.31** |
| stability | V-Dem's 0.08 → the four WGI series: political stability and rule of law 0.14 → **0.16**, control of corruption and government effectiveness 0.10 → **0.12** |
| infrastructure | EGDI's 0.06 → electricity **0.17**, grid **0.15**, LPI **0.15**, renewables **0.13**, internet users **0.10**, airport distance **0.09**; broadband, mobile subs, mobile broadband and road density unchanged (the rounding lands there) |
| community | EF's 0.12 → happiness **0.25**, social support **0.18**, mean years of schooling **0.16**, protected-area access **0.16**, HCI **0.14**, expat density **0.11** |
| fit | EF's 0.34 → `winter_severity` **0.55**, `timezone_overlap_us_hours` **0.45** |
| legal_access | property rights' 0.20 → `foreign_ownership_rules` **0.41**, `visa_pathways` **0.35** |
**Seismic hazard keeps its 0.20 and a coverage flag, and the mechanism is worth stating
because it is the honest version of a bad situation.** There is no open, commercially usable
global seismic hazard grid: GEM's is CC-BY-NC-SA, and what USGS serves openly is a point
service covering **the United States and its territories only** (ISSUE-110). So the indicator
has a value for US units and is missing for the other ~3,550 L1 units. A missing indicator is
never imputed (§4.3): it is dropped from the pillar's active set for that unit, safety
renormalizes over what remains, and the coverage penalty lowers that pillar's confidence
(§5.2). **The failure mode this leaves is real and must not be hidden by the confidence
number alone: a high-hazard non-US city — Tokyo, Santiago, Istanbul — is not penalised for
seismicity at all, so its safety score is optimistic.** Global coverage is a v0.2 item; the
alternatives were licensing GEM (costs money this phase does not have), assembling national
hazard maps one country at a time, or dropping seismic entirely.
**Two rule-7 exposures are tolerated rather than closed, and both are named on §15.2:** `hale`
(WHO GHO, `review`, no mirror) and `tax_burden` (Heritage EFI, `review`, no instrument
published). Both are scored today. The asymmetry with `property_rights_index` is deliberate:
IPRI is `false` — decided against us — so that id could not wait, whereas Heritage is merely
unread. If Heritage comes back blocked, `tax_burden` follows and `legal_access` loses another
0.12.
> *Read this paragraph as dated.* It was written on 2026-08-29/30 and is left as written. `hale`
> was demoted by D10-B at the `0.1.0` freeze, and `tax_burden` by **D12** at `0.1.1` — not
> because Heritage came back blocked, but because *unread is not cleared* and D12 stopped
> waiting for the difference. The 0.12 it warns about did leave, and §16 records where it went.
#### 3.6.5 What the registry actually got when it was staged (2026-08-30, wave B)
§3.6.4 reconciled the registry with the licences. Staging it then reconciled it with the data.
Four of §3.6.1–3.6.4's carriers did not behave as written, and each correction is recorded here
rather than edited into the tables above, because the freeze diff needs the difference between
what was planned and what shipped.
| planned | shipped | why |
|---|---|---|
| `uhc_service_coverage_index` from WDI `SH.UHC.SRVS.CV.XD` | same id, **World Bank database 16** (Health Nutrition and Population Statistics) | the id is **archived in the default WDI database** and answers HTTP 200 with an error body, not a 404. Live and identical in database 16, same CC-BY licence (ISSUE-121) |
| `mean_years_schooling` from "the World Bank's UIS republication" | `UIS.EA.MEAN.1T6.AG25T99`, **EdStats, database 12**, registered as a `wdi_macro` variable | the series is real and CC-BY, but not in the default database; frozen at 2024-06-25, 161 economies, 94 % of world population, never trended (ISSUE-116 closed) |
| `social_support` from OECD Regional Well-Being's Community topic at **TL2** | the **national** How's Life series (`DSD_HSL@DF_HSL_CWB`, `MEASURE=7_1`) | **the TL2 flow does not exist.** All 1,546 public OECD dataflows enumerated; no flow at any territorial level carries a social-support measure, and the regional well-being composite is gone entirely. The national series is the same quantity at a coarser level, which the indicator's own transform already permits. 47 economies, 26 % of world population, every value `inherited_L0` (ISSUE-122) |
| `oecd_regional_wellbeing_index` display-only from OECD TL2 | **no carrier; stays unsourced** | same finding. Left *pending* rather than *retired*: nothing about its licence is the problem (ISSUE-122) |
| `real_house_price_trend_5y` from BIS RPPI | OECD `DSD_AN_HOUSE_PRICES` `MEASURE=RHP` | BIS is `license_ok_commercial: review`, and a **scored** indicator fed only by a `review` carrier is the rule-7 configuration V-Dem was demoted for. Same quantity, cleared licence. **Lost:** ~60 economies → 47, and history from 1927 → 2015. BIS stays registered pending counsel Q20 (ISSUE-128) |
**One thing the wave closed that §3.6 did not expect to close.** ISSUE-010 — open since Phase 0,
where `social_support` was unsourced and only the World Happiness Report's ladder
*decomposition* could be staged — is closed. The community pillar now carries the variable
itself for 47 economies rather than a contribution in ladder points for 153.
**And one it opened.** `conflict_event_density_50km` is implemented exactly as §3.6.2 and its
spec specify — a 50 km buffer around the population-weighted centroid for L0 and L1 — and at
CITY level it works: Khartoum reads 90.8 events/year, Luhansk 1,677. **But `L0:SDN` reads 0.0**,
because Sudan's national population-weighted centroid has no qualifying event within 50 km. On
a `lower_is_better` indicator holding 0.31 of the safety pillar, 0.0 is not a gap the coverage
penalty can discount — it is a positive claim that Sudan is peaceful, and it is worse than the
2008 homicide rate this indicator was added to replace (§3.5.4). **A point buffer is the wrong
construction above CITY level.** The repair — a population-weighted mean of the unit's
children's densities, or events per 100k over the whole unit — is a *different quantity* and
needs this section changed, not the ingester. Filed as ISSUE-124 and on the §15.1 list.
**Closed by §3.6.6.**
#### 3.6.6 ISSUE-124 closed: the conflict indicator is deprecated and replaced (2026-08-30)
`conflict_event_density_50km` is **deprecated**. It is listed in `deprecated_indicators` in
`methodology/v0_1.yaml` with `superseded_by: [conflict_events_per_100k]`, its spec carries
empty `sources` **and** empty `expected_sources` — the enforceable half of "nothing may stage
it" — and `staging.checks` treats a row for it as a problem rather than a note.
**The replacement.** `conflict_events_per_100k` — safety, `lower_is_better`,
`weight_in_pillar` **0.31**, the carrier moved intact so no other weight changed and the
pillar still sums to 1.0. UCDP GED events (all three violence types, `where_prec <= 3`,
trailing 5 complete calendar years), divided by unit population, expressed per 100,000 people
per year.
**Two constructions, one per geometry, stated rather than implied:**
| level | event set | why |
|---|---|---|
| L0, L1 | events whose coordinates fall **inside the unit's compute polygon** (DuckDB spatial `ST_Contains` against `geo_units.geo.parquet`) | a 50 km circle around a national population-weighted centroid is not a country |
| CITY | events within **50 km of the city point**, unchanged | a city here has no administrative polygon (its compute geometry is a 25 km analysis buffer), and "violence near where I would live" is the relocation question |
The two levels are normalized against separate cohorts (§4.1), so a construction that differs
by level does not put two quantities into one ranking.
**Why population and not area is the denominator.** The quantity is exposure to organised
violence. A raw event count makes every large country look violent and every small one safe;
an area denominator makes an empty province with one incident look catastrophic.
`homicide_rate` already uses per-100k in this pillar, so the two safety variables are now read
the same way. **Zero is a measurement, not a gap** — a unit with population and no qualifying
event inside it stages 0.0, exactly as `volcanic_hazard` stages the non-volcanic world. A unit
with no population figure gets no row at all (11 of 7,370): a rate with no denominator is not
a number.
**Why the id had to change rather than be redefined.** "50km" is false for the 3,486 polygon
units under the new construction, and the unit of measure moves from events/year to events per
100k per year. Redefining in place would publish a different quantity under a name already
used for another one — the error §3.4.3 refused for `drought_spei` and §3.6.3 refused for
`food_import_dependence`. CLAUDE.md §4: deprecate, never rename.
**What it reads now**, which is the check that mattered: `L0:SDN` **0.72** events/100k/yr
against Khartoum's **3.69** and El Fasher's **35.9** — same order of magnitude, correct
ordering, no longer a claim that Sudan is peaceful. Globally the worst L0 units are Palestine
18.3, Ukraine 15.3, Lebanon 4.2, Syria 2.8, Mexico 2.1, Afghanistan 1.9, Somalia 1.6, CAR 1.4,
Burkina Faso 1.3, Mali 1.2, Yemen 1.0. 2,041 of 7,359 units carry a non-zero value.
**Declared limitation.** At CITY level the 50 km buffer reaches beyond the city's own
population, so a small city near a large conflict zone reads high — Chula Vista (USA) at 176.9
is picking up Tijuana. That is the intended direction for a relocation question and it is one
more reason the CITY cohort is normalized on its own.
#### 3.6.7 The manual tables are staged, and the categorical→score mappings written down (2026-08-30, ISSUE-134)
`0.1.1` shipped with `legal_access` scoring **nothing for any unit on Earth** (§16.3). The
cause was not the demotion of `tax_burden`; it was that PR #1's [HUMAN]-approved tables had
never been ingested, so 0.86 of the pillar sat on files no scoring run had ever opened. This
closes that. It is a **data patch under `0.1.1`** — new rows under unchanged specs, which §12
calls a patch and gates on the agent — and it moves no weight, threshold, formula or id.
`data/manual/README.md` §6 says the tables store **facts, not scores**, and that the
categorical→0–100 mapping "belongs in `packages/scoring` and `docs/METHODOLOGY.md`, where it is
versioned and can be argued with". None of the three specs written in wave A named the columns
the approved tables actually carry, so the mappings are defined here and mirrored into
`indicators/*.yaml` and `pipelines/staging/transforms/manual_tables.py`.
**`foreign_ownership_rules`** — one enum value to one ordinal, no interpretation added:
| `foreign_ownership_status` | value | |
|---|---|---|
| `freehold` | **4** | same terms as nationals, no approval |
| `freehold_with_conditions` | **3** | freehold, with disclosure or sectoral conditions |
| `leasehold` | **2** | no freehold; lease or trust vehicle carrying full beneficial rights |
| `restricted` | **1** | permit/authorisation regime, or a substantial geographic or sectoral bar |
| `banned` | **0** | barred to non-residents |
| `unknown` | *no row* | "we looked and could not establish it" — a null, never a 0 |
Staged: NOR 4 · PRT 4 · USA 3 · CHE 1 · NZL 1 · MEX 1 · EGY 1 · IND 1.
**The `regional_exceptions` column is deliberately not scored.** Mexico's 50 km coast / 100 km
border restricted zone, Egypt's Sinai ban and the US state farmland statutes are *geographic*,
not national. A country row cannot represent any of them — the [content] agent says so in its
own sidecar — and folding them into a national ordinal would produce a number that is wrong for
both halves of each country. They remain the L1-override gap that
`foreign_ownership_l1_overrides.csv` closes in Phase 1.
**`visa_pathways`** — the scale names four routes (digital-nomad, passive-income, investment,
ancestry), so the value is the **sum of four route credits** over `nomad_visa`, `golden_visa`,
`retirement_visa` and `ancestry_citizenship_pathway`:
| credit | values | reading |
|---|---|---|
| **1.0** | `yes`, `yes_no_real_estate` | a dedicated, published, open route. `yes_no_real_estate` is full credit because only the programme's *property* option closed (Portugal's ARI), and property is `foreign_ownership_rules`' business, not a residence route's |
| **0.5** | `partial`, `no_formal`, `no_formal_programme`, `yes_contested`, `yes_quasi`, `yes_recently_tightened` | the route exists but is materially qualified: permitted without being dedicated (NZ remote work on a visitor visa), no published programme but a route used in practice (MX solvency residence, CH lump-sum taxation), under active legal challenge (US "Gold Card"), not-quite-citizenship (India's OCI), or just re-timed by statute (PT Organic Law 1/2026) |
| **0.0** | `no` | |
| *no row* | any `unknown` among the four | a count cannot be composed out of unknowns: one unresearched route would read as a closed one |
Staged: PRT 3.5 · NZL 2.5 · MEX 2.0 · EGY 2.0 · CHE 1.5 · USA 0.5 · IND 0.5 · NOR 0.0.
**A row may veto itself.** `L0:YEM` and `L0:SDN` carry, inside the approved diff, the sentence
*"EVERY FIELD IN THIS ROW IS LOW CONFIDENCE AND SHOULD NOT FEED A SCORE … Prefer an explicit
'insufficient data' with a visible low-confidence badge (CLAUDE.md §3.6) over any imputed
value."* The sidecar's reviewer checklist put exactly that question to [HUMAN] and
`approval_notes` records no answer, so the row's instruction stands and the ingester honours it
by reading the phrase **out of the table** rather than from a hand-typed exclusion list. If a
refresh drops the phrase, the row scores again and the change is visible in the diff that goes
to the gate.
**`desalination_capacity_per_capita`** — operational capacity only, per 1,000 people. The
compensator of REVIEW §1 Q2.1 exists so an arid coastal unit with real capacity never *reaches*
the veto threshold, which is what makes `plant_status` load-bearing: crediting announced
capacity would suppress a veto that should fire. `operational` counts;
`under_construction`, `standby`, `announced` and `announced_target` count for nothing on the
`now` horizon; `none_identified` with capacity 0 (Arizona) is an **affirmative absence** and
contributes a real zero. Denominator: `population_thousands` from
`data/reference/iso_crosswalk.csv`, so the staged value does not depend on whether the geo
artifact is on disk. Staged: **EGY 15.32 · USA 0.82 · MEX 0.16 · IND 0.14** m³/day per 1,000
people, and nothing else — four countries, so the L0 reference distribution behind the
normalization is **n = 4**, all of which have capacity. Portugal is absent on purpose: its only
asset is under construction until 2028 and may not compensate now.
**What it cost and what it bought.** `legal_access` scores **949 units** (8 L0 + 225 L1 +
716 CITY) instead of none, at confidence **0.688** at L0 — which is `0.86 × 0.80`, the two
staged indicators' weight times the `estimated` quality multiplier every manual row carries.
Access moves by up to **25.4** points and the FCS by up to **2.3** (the desalination
compensator). All 20 anchors pass. The full accounting is in `CHANGELOG.md` under `[0.1.1]`
and in the QA report.
**Two things this does not fix**, recorded so neither reads as an oversight:
* **Coverage is the Phase-0 ten and nothing else.** `legal_access` is null for ~242 of 250
countries and every child of them. That is honest — a country absent from the table is a
country nobody has researched, not a country with no rules — and §7 confidence carries it.
* **`passport_mobility_henley` keeps 0.14 while staging nothing.** See §3.7.
### 3.7 What is still not sourced, and what that costs
Written down because a coverage hole nobody names is a coverage hole nobody fixes, and because
the confidence figure is only honest if the reader can see what is behind it.
| id | pillar · weight | why there is no value | consequence |
|---|---|---|---|
| `desalination_capacity_per_capita` | water · 0.06 | **PARTLY CLOSED 2026-08-30 (§3.6.7).** Four countries staged from the approved manual table (EGY, USA, MEX, IND); GWI DesalData, the only comprehensive register, is still a paid subscription and a [HUMAN] gate | the compensator of REVIEW §1 Q2.1 now works *where the table reaches*. Everywhere else an arid coastal city is still uncompensated and vetoed on water, and the L0 reference distribution is n=4 |
| `rent_yield` | affordability · 0.08 | every global yield dataset is a scraped portal aggregate (barred, CLAUDE.md §3.4) or paid | affordability carries 0.92 of its weight; OECD's price-to-rent ratio is the substitute and would need its own id |
| `passport_mobility_henley` | legal_access · **0.14** (§16.2) | Henley is proprietary. **The id names the source**, so under the ids-describe-their-quantity rule it may only ever be filled from Henley | **the pillar's live gap, and the weight stays.** Not a rule-7 violation: rule 7 bars a *restricted source* from carrying weight and this id has `sources: []` — it is fed by nothing, so `validate_restricted_indicator_weights` has nothing to report and the gate is green. The 0.14 is charged to **confidence**, where §5.2 already puts it (active weight in the denominator, present weight in the numerator), and `pillar_score` runs over the 0.86 that is there. Zeroing it would declare the pillar fully measured at 0.86 of its design and *raise* every unit's confidence for a quantity we do not have. Restoring it needs a real source: deprecate this id, register `visa_free_destinations` |
| `foreign_ownership_rules`, `visa_pathways` | legal_access · 0.46 + 0.40 | **CLOSED 2026-08-30 (§3.6.7).** PR #1 merged; the tables are staged | `legal_access` scores 949 units — the Phase-0 ten's countries and their children — and is null for the rest of the world. Coverage, not absence: the pillar exists again |
| `grid_reliability`, `broadband_speed_fixed`, `mobile_broadband_speed`, `road_density` | infrastructure · 0.14 + 0.06 + 0.04 + 0.06 | Ookla/M-Lab need accounts ([HUMAN]); no free global SAIDI series | infrastructure carries 0.70 of its weight |
| `sovereign_rating` | stability · 0.04 | the three agencies' ratings are proprietary | small |
| `expat_community_density`, `climate_analog_shift_km`, `disaster_loss_index`, `cyclone_track_density`, `wildfire_fwi`, `soil_quality`, `growing_season_length`, `tree_cover_loss_trend`, `insurance_availability`, `construction_cost_index`, `cost_of_living_index`, `wetbulb_warmest_month_proxy`, `protected_area_access` | various | raster or licence work not yet done; `cost_of_living_index`'s natural source (Numbeo) is barred outright | charged to the coverage penalty, never imputed (§4.3). The full weighted-but-silent list is published on `/methodology` (ballot C2) |
**New holes opened by the 2026-08-30 licence reconciliation (§3.6.4).** These are not "not yet
sourced" — they are *sourced and barred*, which is a different and more permanent kind of gap:
| id | pillar · lost weight | why there will be no value | consequence |
|---|---|---|---|
| `agro_suitability_gaez` | food · 0.17 where missing | GAEZ v4 verified CC-BY-NC-SA 2026-08-31 — the catalog's open "verify" closed against us | food has no licence-clean sub-national path; the id waits (the quantity is not source-defined), the weight stays charged to confidence |
| `terrorism_index` | safety · 0.16 | IEP sells a commercial licence and counts display as commercial use | safety measures organised violence (UCDP) rather than terrorism as such; attacks below UCDP's 25-death threshold are invisible |
| `egov_index` | infrastructure · 0.06 | UN Secretariat terms bar commercial use *and* derivative works | no free global measure of digital service delivery remains; connectivity is not administrative capacity |
| `english_proficiency` | fit · 0.34 · community · 0.12 | EF publishes no data file and no licence | **fit is two in-house derivations**; the one indicator measuring day-to-day functioning is gone |
| `property_rights_index` | legal_access · 0.20 | IPRI bars commercial reproduction; Heritage publishes no instrument | Access loses its own expropriation/title reading; `wgi_rule_of_law` carries the signal inside the FCS only |
| `vdem_liberal_democracy` | stability · 0.08 | CC-BY-SA, unresolved | democratic institutions are WGI + Freedom House, and Freedom House is itself `review` |
| `seismic_hazard` outside the US | safety · 0.20 where missing | no open global grid exists | non-US units take the coverage penalty; high-hazard non-US cities are scored optimistically |
**Declared redundancy, retired with the indicator:** the GTI/UCDP overlap this section used to
declare is gone, because GTI is gone. The `egov_index` overlap against `internet_users_pct` and
`human_capital_index` is likewise moot. Neither is a win: in both cases the redundant part is
what survived and the additional information is what was lost.
---
## 4. Normalization
Every indicator becomes a 0–100 *goodness* score by percentile rank **within its
geographic level**, against a **frozen reference distribution**.
### 4.1 The frozen reference distribution (binding, REVIEW §2)
Percentile ranks are zero-sum: they move when the peer set moves. Therefore:
> The reference distribution for `(indicator_id, level, horizon, scenario)` is computed
> **once per methodology version** and frozen. Units added or refreshed inside a version
> are scored **against the frozen distribution**. They do not reshuffle anybody else.
A reference distribution is the tuple
`(indicator_id, level, horizon, scenario, sorted_values, n, p_low, p_high, methodology_version)`.
It is built from the reference cohort's raw values (missing values dropped) with:
- `p_low` = 2nd percentile, `p_high` = 98th percentile, both by **linear
interpolation between the two nearest order statistics** (NumPy's default
`method="linear"`), computed on the *unwinsorized* reference values;
- `sorted_values` = the reference values **after** clipping to `[p_low, p_high]`,
sorted ascending.
Freezing means: the pipeline computes these once for a version and stores them; the
scoring package never rebuilds a reference distribution while scoring.
**Where the frozen library lives.** `data/staged/reference_library/<methodology_version>/`
— `reference_distributions.parquet` (one row per key, carrying `n`, `p_low`, `p_high` and
the clipped sorted values) plus a `manifest.json` with its sha256 and row count. It is
**methodology-version state, not a cache**: every published number under a version is
scored against that file, and re-freezing it is a version event (§12), not a rebuild.
Written by `uv run marts reference-library` (`pipelines/marts`), because
`packages/scoring` is pure and does no I/O.
**A value with no frozen reference is not scored.** If `(indicator, level, horizon,
scenario)` has no frozen distribution, the row is dropped and reported in
`missing_references`. Scoring it against some other cohort's distribution would silently
misrank it, which is worse than a coverage penalty — so it takes the coverage penalty.
**Known limitation (0.1–0.2), closed at `0.3.0` for the physical pillars.** Rank
normalization measures *relative* standing: if every place gets hotter, nobody's climate rank
falls — under `0.2.0` the mean now → 2050 score change of every projected indicator was 0.00
by construction. From `0.3.0` the climate and water pillars' **anchored** indicators are
scored on frozen physical curves (§4.4); the six other pillars stay ranked and the limitation
still applies to them. Say so on `/methodology`.
### 4.2 The formula
For raw value `x` against reference `R`:
1. **Winsorize:** `x' = min(max(x, p_low), p_high)`.
2. **Mid-rank percentile** against the (already clipped) reference values:
```
pct(x') = 100 · ( #{v ∈ R : v < x'} + 0.5 · #{v ∈ R : v == x'} ) / n
```
3. **Direction alignment:**
```
score = pct(x') if direction is higher_is_better
score = 100 − pct(x') if direction is lower_is_better
```
Result is always in `[0, 100]`.
Mid-rank (rather than `#{v ≤ x}/n`) is used so that ties are not rewarded for being
tied and so the top and bottom of the cohort are not forced to exactly 100 / 0 — which
would put a unit at the clamp boundary of §5.3 for purely ordinal reasons. Consequences,
all intended and all tested:
- `n = 1` → every value scores 50.
- All reference values identical → every value scores 50 (an indicator with no variance
carries no information, and 50 says exactly that).
- A value above `p_high` and the winsorized top of the cohort receive the same score.
### 4.3 Missing values
A missing indicator is **not** imputed and **not** treated as zero. It is dropped from
the pillar mean and charged to the pillar's coverage penalty (§5.2). Imputation would
manufacture confidence we do not have.
---
### 4.4 Absolute anchoring (binding, `0.3.0` ballot A1–A5)
An **anchored** indicator's score is read off a frozen piecewise-linear **value → score
curve**: linear between breakpoints, flat beyond the ends (never extrapolated), `nan` for a
missing value (§4.3). Only indicators of the horizon-eligible pillars may be anchored; the
config refuses a curve anywhere else. The same raw value scores the same at every level,
horizon and scenario.
**What `provenance` claims — and what it does not.** Every curve publishes a `provenance`
(`cited | methodology_choice | unverified`) and a prose `basis`, and both are published on
`/methodology`. `provenance` grades **where the curve's class boundaries come from**:
`cited` when the boundaries that separate one published class from the next are a source's
own, `methodology_choice` when at least one boundary is ours because nobody bands the
quantity, `unverified` when a citation is claimed but not yet checked. It does **not** claim
that the *scores* are cited, and on nine of the eleven curves they are not. A published
threshold and the number of points that threshold costs are two different decisions, and
only the first is ever somebody else's: Aqueduct says withdrawals at 40 % of renewable
supply is "high" stress, but that this should score 45 rather than 40 or 50 is our
judgement. The two curves whose scores are not ours either are
`safely_managed_drinking_water_pct` (the score *is* the published percentage) and
`drought_risk_aqueduct` (Aqueduct's own 0–5 class score, linearly rescaled). Every
`methodology_choice` `basis` names which of the two decisions is ours; read it before
quoting the label.
**How much of this is our judgement.** Six of the eleven curves are `methodology_choice`,
and **all six sit in the climate pillar**: `tasmax_warmest_month` (0.16),
`coastal_flood_slr_exposure` (0.14), `riverine_flood_100yr_exposure` (0.12), `wildfire_fwi`
(0.10), `heat_degree_months_30c` (0.08) and `cyclone_wind_100yr_ms` (0.06). Together they
carry **0.66 of the climate pillar's 1.00** — two thirds of the pillar — out of the 1.48 of
in-pillar weight anchoring touches at all, so **45 %** of anchored in-pillar weight. Weighted
the way the composite actually feels them (climate 25, water 18): 25 × 0.66 = **16.5** of the
25 × 0.78 + 18 × 0.70 = **32.1** anchored composite points, i.e. **51 %** of anchored
composite weight and 16.5 % of the default FCS. Both figures are true and they differ only
in denominator, which is why both are printed here rather than whichever is smaller. Water is
the other side of the ledger: all four anchored water indicators, 0.70 of that pillar, are
`cited`. These numbers are asserted against the config by
`packages/scoring/tests/test_anchoring.py::test_4_4_publishes_the_real_methodology_choice_share`,
so they cannot rot away from the curve table.
That is a disclosure, not a confession. The alternative to a stated curve is not an objective
one — it is the percentile rank of §4.1, which makes the very same judgement (that 35 °C is
worth whatever the cohort's shape says it is worth) and publishes none of it. A frozen,
versioned, `provenance`-labelled curve can be argued with line by line, and moving one costs
a major version bump and a [HUMAN] gate. A rank cannot be argued with, because it never says
what it thinks 35 °C is worth.
**Freeze rule.** The curve table is versioned state: frozen with the version
(`anchoring.frozen_on`), recorded in the CHANGELOG with its citations, never edited inside a
version. Changing a breakpoint, adding or removing a curve, or moving an indicator between
ranked and anchored is a normalization change — **major bump, [HUMAN]** (§12). An anchored
indicator has **no reference distribution**: the library omits it and QA's rebuild check
(`stale_reference_keys`) fails on one.
**Level.** The anchored curves read above the world's cohort medians, so at `0.3.0` `now`
climate scores rose ~15 points and water ~9 on average with rank correlations ≥ 0.956 at
every level; this was chosen (ballot A3) and disclosed, and a score of 50 now names a
physical condition (a 35 °C warmest month; Aqueduct medium-high stress), not the middle of a
cohort. The level shift is a property of the whole anchored set, cited and chosen curves
alike — it is not evidence that the thresholds are external.
**Water stress at `now`.** `projected_water_stress` carries the Aqueduct baseline at `now`
(`quality_flag: baseline_as_now`) so both horizons score the same indicator set (ballot A5);
C2.4's "no `now` value by construction" is retired.
| indicator | dir | breakpoints (raw → score) | provenance |
|---|---|---|---|
| `tasmax_warmest_month` °C | lower | 24→100 · 30→80 · 35→50 · 40→20 · 45→5 · 48→0 | methodology_choice (TX35 / TX40 cited; 24 / 30 / 45 / 48 and every score ours) |
| `heat_degree_months_30c` | lower | 0→100 · 6→85 · 18→60 · 36→35 · 72→10 · 100→0 | methodology_choice (our quantity; no external band exists) |
| `coastal_flood_slr_exposure` | lower | 0→100 · 0.05→85 · 0.20→60 · 0.50→30 · 1.0→0 | methodology_choice (LECZ defines the quantity, nobody bands it) |
| `riverine_flood_100yr_exposure` | lower | 0→100 · 0.02→85 · 0.10→60 · 0.30→30 · 0.60→10 · 1.0→0 | methodology_choice (ballot C3's set, confirmed at the freeze by F1 against the staged distribution; unprotected hazard) |
| `cyclone_wind_100yr_ms` | lower | 0→100 · 15.39→90 · 28.97→70 · 37.58→50 · 43.46→30 · 51.15→10 · 62.02→0 | methodology_choice (Saffir-Simpson floors, our conversion and scores) |
| `wildfire_fwi` (days FWI ≥ 38) | lower | 0→100 · 6.5→85 · 57→60 · 151→30 · 219→15 · 296→5 · 366→0 | methodology_choice (EFFIS defines the ≥ 38 day; the day counts are ours — pinned to the staged 2005–2024 quantiles by F5; fire weather, not fire occurrence) |
| `drought_risk_aqueduct` | lower | 0→100 · 0.2→80 · 0.4→60 · 0.6→40 · 0.8→20 · 1.0→0 | cited (Aqueduct 4.0 cut points *and* its own 0–5 class score) |
| `baseline_water_stress` = `projected_water_stress` | lower | 0→100 · 0.10→90 · 0.20→70 · 0.40→45 · 0.80→20 · 1.0→10 · 2.0→0 | cited (Aqueduct 4.0 bands; 1.0 / 2.0 and every score ours) |
| `renewable_freshwater_per_capita` | higher | 250→0 · 500→10 · 1000→35 · 1700→65 · 5000→95 · 10000→100 | cited (Falkenmark 500 / 1000 / 1700; 250 / 5000 / 10000 and every score ours) |
| `safely_managed_drinking_water_pct` | higher | 0→0 · 100→100 | cited (JMP / SDG 6.1.1 — the score *is* the percentage) |
The table above is the human-readable mirror of `methodology/v0_3.yaml → anchoring`; the YAML
is normative, and `test_methodology_4_4_table_mirrors_the_yaml` parses this table and asserts
every id, direction, breakpoint and `provenance` matches it. Until 2026-09-03 this paragraph
claimed such a test existed when none did, and in the gap the table drifted: it printed a
curve the config does not have (`riverine_flood_exposure`, retired by C3), omitted one it
does (`cyclone_wind_100yr_ms`), and left `riverine_flood_100yr_exposure`'s breakpoints
unwritten for weeks after they were calibrated. Claiming a test is worse than the
overstatement the claim is meant to reassure about; the test is the fix.
**Open at the freeze — one label, no scores.** Three `cited` curves take the generous reading
of their own label. `baseline_water_stress` / `projected_water_stress` and
`renewable_freshwater_per_capita` draw their class boundaries from Aqueduct and Falkenmark,
but each extends past the last published boundary on breakpoints of our own (1.0 and 2.0;
250, 5,000 and 10,000), and none of the three takes its scores from the source.
`cyclone_wind_100yr_ms` applies the stricter reading to itself — its boundaries *are* the
Saffir-Simpson floors and it still declares `methodology_choice`, because the unit conversion
and the scores are ours. Applied uniformly, the strict reading leaves only
`drought_risk_aqueduct` and `safely_managed_drinking_water_pct` as `cited` and makes the
share above **82 %** rather than 45 %. Which reading the labels take is a freeze-ballot line
for the curve table's owner, not an editorial choice: **no breakpoint and no score moves
either way**, and both readings are stated here so a reader is not relying on the label alone.
**One curve is not inherited.** Ballot A1's table was written against the `0.2.0` indicator
set, where `riverine_flood_exposure` is Aqueduct's *annual expected* exposure share. Ballot C3
replaces it with `riverine_flood_100yr_exposure`, the share of a unit's population inside the
100-year inundation footprint — **a different quantity on a different scale**, so the old
breakpoints do not transfer and were not transferred. **Its breakpoints are the set ballot C3
named (2026-09-02), confirmed by freeze ballot F1 (2026-09-09,
`docs/METHODOLOGY_v0.3_FREEZE_BALLOT.md`)** against the post-5 cm-floor staged distribution
(n = 7,139; median 0.0768, p75 0.1720, p90 0.3210, p95 0.4352): the global median scores 67,
p90 29, p95 21. The curve is shaped like the coastal LECZ curve but deliberately tighter per
share, because a modelled ≥ 5 cm 100-year footprint is not an elevation band. A 2026-09-02
re-fit to the pre-floor distribution (median 0.088, p90 0.362; ISSUE-150) sat in the config and
in this table until the freeze and was never signed; F1 chose not to recalibrate beyond what C3
wrote, and the level shift it records (mean −7.6 against the unsigned set, ordering identical
by construction) is in the ballot. The estimator is the **median of the five-GCM ensemble**;
where that ensemble disagrees about whether a place floods at all, the row carries
`ensemble_divergent` and is charged in confidence, never in the value (F2, §5.2). **The quantity
is unprotected hazard** — Aqueduct models no dikes or levees, so a defended delta scores like an
undefended one; a protection-adjusted variant is under measurement
(`docs/research/2026-09-09_flood_protection_prototype.md`), and any change it earns is a
quantity change on this line brought back to [HUMAN] with numbers. It carries provenance
`methodology_choice` because nobody publishes population-share bands for a 100-year footprint,
and is disclosed on `/methodology` with the rest. The deprecated id carries **no** curve: it
is out of the model, and `0.2.0` stays reproducible without one because `0.2.0` anchors
nothing at all — it percentile-ranks every indicator from `reference_library/0.2.0/`.
**One curve was pinned to data at the freeze.** `wildfire_fwi`'s day counts were written before
any row was staged (its own basis said so), and the first staging — CEMS
`cems-fire-historical-v1`, mean annual days FWI ≥ 38 over 2005–2024, 7,312 units; median 6.5 d,
p75 56.8, p90 150.8, p95 219.1, p99 295.5, max 363.2 — showed them scoring the median unit 86.9
and putting 24.2 % of units below 30 and 6.0 % at exactly 0, the Mediterranean belt at 14–25.
Freeze ballot **F5** re-pinned the counts to those quantiles (median → 85, p75 → 60, p90 → 30,
p95 → 15, p99 → 5, 366 days → 0; the observed maximum, 363.2 d, scores 0.2): the share below 30 falls to 10.0 %, no unit sits
at a hard 0, and the deserts stay at the bottom — which is what fire weather says about them.
**The quantity is fire weather, not fire occurrence:** a fuel-free desert scores low because
its fire weather is severe, not because it burns. The median unit sits at 85 by construction;
whether that is the right centre for a fire-weather day count is queued as **M28**, a 0.3.x
ballot line once 0.3.0 is scored.
## 5. Pillar scores
### 5.1 Active indicator set
For a given persona `P`, the **active** indicator set of a pillar is:
```
active(pillar, P) = { i ∈ pillar : not i.display_only
and (i.gated_by is None or i.gated_by ⊆ P.enabled_gates) }
```
`display_only` indicators (FSI, GPI) are never active. Gated indicators are active only
under a persona that enables their gate.
### 5.2 Score and coverage penalty
Let `A ⊆ active(pillar, P)` be the indicators that have a value for this unit/horizon/
scenario, `w_i` the `weight_in_pillar`, `s_i` the normalized score from §4.
```
pillar_score = Σ_{i∈A} w_i · s_i / Σ_{i∈A} w_i
pillar_confidence = Σ_{i∈A} w_i · q_i / Σ_{i∈active} w_i
```
`pillar_score` is a **weighted arithmetic** mean — inside a pillar, indicators *are*
substitutable and compensators (desalination) are supposed to compensate.
`pillar_confidence ∈ [0, 1]` is the coverage penalty of REVIEW §2 / brief §3.3.2,
extended by a **data-quality multiplier** `q_i ∈ [0, 1]` keyed on the
`quality_flag` column of `indicator_values` (architecture §3.3 requires inherited values
to feed confidence):
| `quality_flag` | `q` | Meaning |
|---|---|---|
| `ok` (or null) | 1.0 | measured at this unit's own level |
| `modeled` | 0.85 | modeled/downscaled to this unit |
| `estimated` | 0.80 | estimated or imputed upstream |
| `inherited_TL2` | 0.70 | inherited from the parent region's regional-statistics value |
| `inherited_L1` | 0.70 | inherited from the parent admin-1 |
| `inherited_L0` | 0.50 | inherited from the country |
| `ensemble_divergent` | 0.60 | the row's own multi-model ensemble disagrees about whether the hazard is present at all — for `riverine_flood_100yr_exposure`, ≥ 1 of the five GCMs < 0.01 **and** ≥ 1 > 0.10 (`gcm_dry=<n>;gcm_wet=<m>` beside it); the value stays the ensemble median (freeze ballot F2, ISSUE-149) |
Every row additionally carries a `resolution_flag` —
`measured_at_unit | measured_TL2 | inherited_TL2 | inherited_L0` — the first-class record
of where the value was measured. A city inherits its region's value before its country's
wherever the region has one. **Depth raises confidence, never rank:** percentile ranks stay
within each indicator's own coverage set, and coverage asymmetry is charged to confidence —
the `seismic_hazard` precedent, generalized. The three inheritance multipliers were
calibrated by `sensitivity.py` before the 0.2.0 freeze (result in the CHANGELOG).
Unknown flags are treated as `estimated` (0.80) — never as 1.0. A composite flag
(`;`-joined tokens, `key=value` read by key) takes the **smallest** multiplier among its known
tokens; none known → 0.80 (freeze ballot F3, ISSUE-154 — until `0.3.0` the whole string was
looked up, so every composite fell to 0.80 whatever it said; an empty flag is `ok`).
`measured_at_unit` and `measured_TL2` (the unit *is* the matched region) carry `q = 1.0`, the
existing `ok` semantics. Where a row carries both a `quality_flag` and a `resolution_flag`, the effective
multiplier is the **smaller** of the two — one epistemic charge, never a double-penalty. The
quality multiplier affects **confidence only**, never the score: an inherited value is still
our best estimate of the value, we are just less sure it applies here.
If `A` is empty the pillar has **no score** (null, not zero) and
`pillar_confidence = 0`.
---
## 6. Composite (FCS)
### 6.1 Weighted geometric mean with pre-log clamp
Let `S` be the set of pillars that have a score, `W_p` the persona's pillar weight.
```
c_p = min(max(pillar_score_p, 5), 100) # clamp BEFORE the log
FCS_raw = exp( Σ_{p∈S} W_p · ln(c_p) / Σ_{p∈S} W_p )
```
The clamp to `[5, 100]` (REVIEW §1 Q2.3) exists so a 0 pillar cannot send the composite
to `−∞`; a place with a 5-level pillar is already devastated by the geometric mean
without numerical drama. Pillars with no score are dropped and the remaining weights
renormalize — the alternative (treating "unknown" as 5) would let missing data look like
catastrophe. Missing pillars are charged to confidence instead (§7).
If **no** pillar has a score, the FCS is null.
### 6.2 Veto
```
veto_flags = [ p for p in S if pillar_score_p < 20.0 ] # strict <, pre-clamp
FCS = min(FCS_raw, 40) if veto_flags and veto_cap_enabled
FCS_raw otherwise
```
- The threshold test is on the **unclamped** pillar score and is **strict**: 19.99
vetoes, 20.00 does not. (This exact boundary is a test.)
- "After compensators" (REVIEW §1 Q2.2) is satisfied structurally: compensators are
indicators inside pillars (§3.2), so the pillar score the veto sees is already
post-compensation. There is no separate compensation step.
- **Pro may disable the cap** (`veto_cap_enabled = false`). **The flag is always
returned and always displayed** — disabling the cap changes the number, never the
disclosure. A red flag names the pillar: "Water: critical".
> **Known interaction — the veto is a relative threshold for the six ranked pillars.** For
> climate and water it is an absolute one from `0.3.0` (§4.4): a water pillar below 20
> requires its anchored members to read Aqueduct "extremely high" stress and Falkenmark
> absolute scarcity together; a climate pillar below 20 requires a near-40 °C warmest month,
> a long hot season and coastal or flood exposure. Climate-pillar vetoes fell from 6.8 % of
> countries to 0.0 % at `0.3.0` for this reason and no other. The text that follows records
> 0.1–0.2 behaviour and still governs the ranked pillars.
>
> *(0.1–0.2 text, retained.)* Pillar scores are
> percentile ranks (§4), so "below 20" means "in the bottom fifth of its cohort", not
> "below an absolute standard". In a **small cohort the veto therefore fires by
> construction**: with 10 countries, the last-placed country scores ~5 on every pillar
> it is last in, whatever its absolute condition. This is acceptable at Phase 0 (where
> the bottom of the ten is Yemen and Sudan, and the flag is true) but it is *not* a
> property to rely on, and it is the second reason — after §4.1 — that absolute
> anchoring for the physical pillars is the priority for v0.2. Until then: never present
> a veto flag from a cohort of fewer than ~30 units without the confidence chip beside
> it, and QA should treat a veto rate above ~10 % of a level as a smell.
**That smell test is now measured rather than remembered.** `diagnostics.veto_rate_smell`
(0.10) and `diagnostics.veto_rate_tripwire` (0.60) live in the methodology config, the per-level
rate is computed by the engine (`futurecast_scoring.diagnostics`, §11.1), and it is printed by
`uv run marts score`, quoted in `docs/QA_REPORT.md` and asserted by QA — all from one
implementation. The tripwire is a **regression ceiling, never a target**: at Phase 0 the rate is
40–55 % per level for the coverage reason diagnosed in ISSUE-018 (a pillar scored from one
percentile-ranked indicator puts ~20 % of any cohort under the strict-20 threshold *by
construction*, and Phase 0 had four such pillars). The fix is absolute anchoring at v0.2 plus
coverage, not a looser bound.
### 6.3 Quadrant
The two axes compress to the chip of brief §1:
| | Access < 60 | Access ≥ 60 |
|---|---|---|
| **FCS ≥ 60** | `watchlist` | `move_buy` |
| **FCS < 60** | `avoid` | `trap` |
Threshold 60 on both axes; null on either axis → quadrant is null.
The ids above are the contract (data, `/api/v1`, MCP payloads, URL params) and never move.
What a reader sees is the label:
| id | Label | Gloss |
|---|---|---|
| `move_buy` (`move` on the wire) | **Move / Buy** | Holds up, and you can realistically get in. |
| `watchlist` | **Wait** | Holds up, but cost or legal access is the barrier — wait for a window. |
| `trap` | **Trap** | Easy and cheap to enter, for reasons that show up in the score. |
| `avoid` | **Avoid** | Neither resilient nor accessible. |
*2026-09-02 — label only:* `watchlist` is shown as **Wait** (UX audit U-11, approved by
[HUMAN]); it was "Watchlist", which collided with the saved-places feature's own verb on the
same card. Ids, thresholds and the gloss are unchanged, and a display word is not part of the
frozen method, so this carries no methodology version bump. `/api/v1/places/{slug}` and MCP
`get_place` carry the word as an additive `quadrant_label` beside the id.
---
## 7. Confidence
Displayed confidence (REVIEW §2 — `min()` is too harsh as a headline, because a thin
Food pillar at 5 % weight would tank an otherwise well-covered unit):
```
coverage = Σ_p W_p · pillar_confidence_p / Σ_p W_p # over ALL pillars in the
# profile; missing pillar
# contributes 0
confidence = coverage · recency_factor
```
**The minimum is not discarded, it is surfaced separately.** Any pillar with
`pillar_confidence < 0.5` raises a **"thin data: {pillar}"** chip, listed alongside the
score. Both the aggregate and the thin-data list are always shown. Never hide either.
**Recency factor.** For each contributing indicator `i` with an `as_of` date, let
`age_i` be its age in years at the scoring reference date (`(ref − as_of).days / 365.25`,
negative ages clamped to 0):
```
r_i = 1.0 if age_i ≤ grace
= max( floor, 1 − (age_i − grace) / decay ) otherwise
grace = 2.0 years, decay = 8.0 years, floor = 0.5
```
The unit's `recency_factor` is the weight-weighted mean of `r_i` over all contributing
indicators, weighting each by `W_pillar(i) · w_i`. Indicators with no `as_of` are treated
as `r = 1.0` (absence of a date is a pipeline gap, charged to coverage, not twice).
If no indicator carries a date, `recency_factor = 1.0`.
**The scoring reference date is an input, not the clock.** It defaults to the newest
`as_of` present in the input frame, never to `today()` — otherwise the same inputs would
produce different confidences on different days, and scoring would not be reproducible.
A negative age (an `as_of` in the future) is clamped to 0: that is a pipeline bug, not a
freshness bonus.
Confidence is a fraction in `[0, 1]`; it is computed unrounded and **displayed** to two
significant figures.
### 7.1 The published-rankings floor (`rankable`)
*Added at the `0.1.0` freeze by REVIEW **D10-C** (Eric, 2026-08-30).*
```
rankable = (fcs is not null) AND (coverage ≥ 0.40)
```
**Measured on `coverage`, not on `confidence`.** The two differ by the recency factor, and a
place whose pillars are all present but eight years old has `confidence = 0.5` with
`coverage = 1.0`. That place belongs in a ranking with a freshness chip; dropping it would be
punishing staleness twice, once in the number it already discloses and once by removing the
place. The floor is a statement about **how much of the model actually ran**, nothing else.
The engine already computes `coverage` for every unit (§7, above) — `rankable` adds no new
measure, only a threshold on an existing one.
**What `rankable = false` does and does not mean.** It does **not** un-score the unit. The
score stands, the Place Card renders, ⌘K finds it, the globe colours it, and the API returns
it. What it forbids is one specific claim: **a unit below the floor may not appear in a
published top-N list, ranking table, "best places" surface or league position.** Where a rank
would go, the surface reads *"insufficient data to rank"* and the reason is one click away.
**Why a floor at all, and why 0.40.** A percentile rank is a comparative claim, and a unit
scored from a third of its weight is making a much weaker claim that looks identical to a
strong one. Réunion sat in the global L0 top ten at coverage **0.16** — a real number, an
indefensible rank. 0.40 is "at least two fifths of the model ran"; it was chosen by the
[HUMAN] ballot, not tuned to produce a particular list, and it was not moved afterwards to
rescue or remove any specific place. Changing it changes who appears in a published ranking,
so it is methodology state (§12) and a **[HUMAN]** gate.
**Where it is emitted.** `composite_scores.rankable` (per unit × horizon × scenario ×
persona — coverage is weight-weighted, so the flag is persona-dependent), `units.json`, and
the tile as an integer `0/1`. It is written `true` **and** `false`, never omitted: "below the
floor" is the finding, and an absent flag would be indistinguishable from a unit that
qualifies. A unit with no composite row at all has no flag and is likewise not rankable.
---
## 8. Horizons
**Binding rule (CLAUDE.md §3.5 as amended by the `0.3.0` ballot, REVIEW §2): an indicator
carries `2035` / `2050` values only if its spec declares a `projection` block —
`family: climate` (climate, water, projected hazards, crop yield; scenarios `ssp245` /
`ssp585`) or `family: demographic` (UN WPP medium variant; scenario `wpp_medium`). Stability,
safety, institutions, community and every other social quantity are never projected; a
projected row for an indicator without a `projection` block is a pipeline violation.**
Resolution for a requested `(horizon, scenario)`:
| Rows | Used |
|---|---|
| indicator with `projection.family: climate` | `horizon = requested`, `scenario = requested` (falling back to `now`/`obs` where that indicator has no projection, flagged `no_projection`) |
| indicator with `projection.family: demographic` | `horizon = requested`, `scenario = wpp_medium`, regardless of the requested SSP — the UN medium variant is not an SSP and is never relabelled as one |
| every other indicator | **always** `horizon = now`, `scenario = obs` |
Eligibility is per **indicator**, not per pillar (ballot C4): health carries one projected
member (`dependency_ratio`) and seven `now` members; food carries one
(`crop_yield_change_2050_pct`). A future-horizon row for an indicator with no `projection`
block is a **pipeline error**: the resolver drops it and records it in the returned
violations list. It never enters a score.
**Input contract.** `indicator_values` is keyed on
`(unit_id, indicator_id, horizon, scenario)` — **one value per key**. A duplicate key is
a pipeline bug; the resolver keeps one row deterministically (stable sort on
`unit_id, indicator_id, horizon, scenario, value`, first wins) and reports the count as
`duplicate_indicator_rows`, so that scoring can never depend on the row order of the
input. An indicator shared by both axes (`english_proficiency`, display-only since
§3.6.4 but still the shape the resolver must handle) carries exactly one value per unit —
the axes read the same number, they do not each supply one.
Therefore:
```
FCS(2050, ssp245) = geometric mean of [ projected climate, projected water,
partly-projected health and food,
CURRENT stability, infrastructure,
safety, community ]
```
**Labeling is part of the methodology, not the copy.** An indicator's horizon chip is its
family's label — **"2050 climate horizon (SSP2-4.5)"**, **"2050 demographic horizon (UN
medium)"** — and the composite's chip names every family present: **"2050 horizon — climate
SSP2-4.5 · demographics UN medium"**; never "forecast", "2050 score" or "predicted".
Indicators with no projection carry a 5-year trend arrow (`trend_5y`), never a projection.
**Raw physical change is displayed, always (ballot E1).** Every 2050 view states each
projected indicator's change in its own unit beside the score — "hot months above 30 °C:
3 → 8", "warmest month +1.4 °C", "water stress 0.66 → 0.78" — with the cohort median in
muted ink. A surface without it is non-conformant.
**Local vs global warming (ballot B1).** Every 2050 view states **"+{local} °C here vs
+{global} °C global ({scenario})"** from `local_warming_delta_c` and the global mean of the
same ensemble and window; the global figure is never taken from another report's table.
**Years between the modelled horizons are interpolated, and say so.** Three years are
modelled — `now`, `2035`, `2050` — and a surface that walks a unit through the years
between them (the globe's play control) shows **interpolated frames**. A frame is built by
interpolating the **raw value** of each *anchored* indicator (§4.4) between its two
bracketing modelled anchors and re-deriving the score through the frozen curve. **A score
is never interpolated.** An indicator that is projected but still rank-normalized is never
interpolated at all — §4.1 freezes its reference distribution per horizon, and there is no
frozen reference for a year we do not model — so it carries its bracketing horizon's
modelled value and changes once, at the boundary. An indicator with no `projection` block
is its `now` value in every frame. Every emitted value records which of these it is
(`modelled` · `interpolated` · `held_modelled` · `held_now`); a frame that reuses a `now`
value and a frame that says it reuses a `now` value are the same frame.
The interpolation runs on **global-mean warming, not calendar time**: the abscissa is the
land-mean ΔT trajectory of the same ensemble and windows that produced the anchors, because
the threshold quantities the physical pillars score (days above a temperature,
degree-months, a stress ratio) are strongly nonlinear in years and much closer to linear in
ΔT. The `demographic` family interpolates on calendar time and never on an SSP. Both
choices were measured by holding the modelled 2035 frame out and predicting it from `now`
and `2050` (`docs/research/2026-09-03_horizon_frames.md` §2).
**Two guarantees and one disclosed departure.** A frame at **2050** is the modelled 2050
frame for every indicator — playing an animation to its end lands on the published 2050
score, never near it. A frame at **2035** carries the modelled 2035 value of every
indicator that has a 2035 row. An anchored indicator that has a 2050 row and **no** 2035
row is interpolated *through* 2035 rather than held at `now`, so at 2035 the frame departs
from the published 2035 horizon for that indicator; the departure is published per
indicator in score points beside the frames.
*As of 2026-09-03 exactly one indicator is in that set.* Queue item M26 staged Aqueduct's
**2030** horizon for `riverine_flood_100yr_exposure` (read as 2035 and flagged
`horizon_offset_2030`, alongside the existing `rcp4p5_as_ssp245`), which removes its departure
**structurally** rather than shrinking it: with a 2035 row present the frame and the published
horizon read the same value and the indicator leaves the departures table entirely. It was not
a small error being tidied away — the interpolation had been standing in for a mean of 1.18
(SSP2-4.5) / 1.13 (SSP5-8.5) score points, p95 4.7, max 84.6. **`cyclone_wind_100yr_ms` remains
and is the larger of the two** — mean 2.54, p95 9.6, max 30.1 — because STORM publishes no
mid-century frame at all: the deposit defines `present = 1979-2014` and `future = 2015-2050`,
one delta per GCM over IBTrACS, with no epoch dimension to select. So this halved the number of
departing indicators without touching the size of the worst one, and the honest reading of this
paragraph is that it now describes cyclone alone.
**Labeling.** An interpolated year never gets a horizon chip of its own. A surface showing
one states the year as **interpolated between the modelled horizons** — "2042 · between the
2035 and 2050 climate horizons (SSP2-4.5), interpolated" — and says once, in reach of the
control, that only `now`, 2035 and 2050 are modelled. **"2042 forecast", "2042 score",
"projected to 2042" and "in 2042" are banned exactly as "2050 forecast" is**: a frame says
where the modelled trajectory passes, not what a year will be like. A surface that animates
scores without the interpolated label is non-conformant, exactly as one that drops the raw
physical change is.
The tile schema follows the pillars that can move: 8 now-pillars + (the horizon-carrying
pillars × 2 horizons × 2 scenarios) attributes (REVIEW §2), not 48.
---
## 9. Access Score
Same normalization machinery (§4), same coverage/quality penalty (§5.2), same recency
(§7). Differences:
```
access = Σ_p W_p · pillar_score_p / Σ_p W_p # ARITHMETIC (see §1.1)
```
- **No clamp, no log, no veto, no cap.** Access has no veto because there is no access
failure that is unrecoverable in the way a water failure is: a banned market is a
closed door, not a dead place, and the `legal_access` pillar score already says so.
- Access is `now`-only. No Access indicator is horizon-eligible.
- Confidence is computed exactly as §7 over the three Access pillars.
### 9.1 Outlook — context beside the score, never inside it
*(Added at `0.3.0` by ballot D1.)*
An **Outlook** panel carries named forecasters' forward-looking economic context for a
country: IMF WEO 5-year GDP-per-capita growth (`review` carrier — free-tier display only,
marked, until counsel Q21), OECD long-term baseline GDP per capita 2050 vs today (licence
`unverified` until OECD's terms are read), World Bank fossil rents % GDP with a
transition-exposure band (< 2 % low · 2–10 % moderate · > 10 % high; CC-BY 4.0), and —
counsel permitting — IIASA SSP2 GDP per capita 2050. Every line shows `(source, vintage,
horizon, confidence)` and the sentence "Outlook is context from the named forecaster, not
part of the score." Outlook carries zero weight in either axis, is never normalized or
ranked, and is never coloured on the score ramp. It is **not a third displayed number**
(ballot D1).
The Pro persona `momentum` (gate `trend_aware`) is the only place measured five-year trends
(`*_trend_5y`) carry weight, additively at 0.05 each in their pillars (§3.1 rule 5); it
prices direction of travel on observed data and projects nothing.
---
## 10. Personas
A persona is a named pillar-weight vector plus a set of enabled indicator gates. Users
may fork and edit weights (Pro); the vector travels in the URL. Validation, enforced at
load and on every user-supplied vector:
1. keys are exactly the eight FCS pillar ids;
2. every weight is finite and `≥ 0`;
3. weights sum to **100** (tolerance `1e-6`);
4. at least one weight is `> 0`;
5. enabled gates are known gate names.
| Persona | Cl | Wa | St | He | In | Sa | Co | Fo | Gates |
|---|---|---|---|---|---|---|---|---|---|
| `default` | 25 | 18 | 18 | 10 | 10 | 6 | 8 | 5 | — |
| `family_relocator` | 20 | 15 | 15 | 16 | 8 | 10 | 12 | 4 | — |
| `retiree` | 22 | 15 | 16 | 22 | 9 | 6 | 8 | 2 | — |
| `nomad` | 18 | 12 | 16 | 10 | 22 | 6 | 12 | 4 | — |
| `land_re_investor` | 24 | 22 | 22 | 6 | 10 | 8 | 4 | 4 | — |
| `resilience_first` | 22 | 24 | 12 | 6 | 8 | 16 | 2 | 10 | `personal_defense` |
| `temperate_steady` | 24 | 14 | 20 | 8 | 8 | 10 | 10 | 6 | — |
| `institutional` | 25 | 18 | 18 | 10 | 10 | 6 | 8 | 5 | — |
`resilience_first` is the **only** profile that enables the `personal_defense` gate, and
therefore the only one under which `firearm_legality` and
`distance_to_active_conflict_km` exist at all (REVIEW §1 Q5). `institutional` shares the
default vector and differs in presentation only (spread/confidence intervals shown).
### 10.1 Filter presets, and `temperate_steady`
**A persona is a weight vector *and*, optionally, a filter preset.** Some of what a persona
means is an emphasis — that is a weight — and some of it is a threshold, which a weight cannot
express: "elevation safe from sea-level rise" is not "weight climate higher", it is "exclude
places above a coastal-exposure line". Presets therefore live in the methodology config, not in
a UI component, because the globe, the rankings table and any export must apply the same
thresholds or they are describing different sets of places under one name.
A clause is `(kind, target, op, value|low+high, horizon, scenario, rationale)`. `pillar` and
`indicator` clauses are always on the **0–100 normalized score** (§4), never on raw units, so
one clause means the same thing for a temperature and for a corruption index; `unit_attribute`
clauses (`pop`, `area_km2`, `level`) are on the raw attribute. Every clause carries its
`rationale` — a preset whose clauses cannot say why they are there is a preset nobody can audit.
Validation at load: targets must exist, and a clause at a projected horizon must name a
climate or water quantity (CLAUDE.md §3 rule 5).
**`temperate_steady`** implements REVIEW addendum 2026-08-29-b **D5** (Eric's own profile: *mid-size
polity; elevation safe from SLR; strong social-democratic institutions; long growing season;
hunting access; low natural-disaster exposure; moderate, reliable rain and sun; low
seasonality*). The weights carry the emphases — Stability 20 for the institutions clause, Safety
10 because natural hazard lives in Safety (§3.1 rule 2) and "low natural-disaster exposure" is a
Safety statement here rather than a personal-defence one, Food 6 for the growing season. The
preset carries the thresholds: population 1–20 M; Stability ≥ 70; Climate ≥ 60 **on the 2050
horizon**, not only today; Water ≥ 50; coastal-exposure, precipitation-variability,
growing-season, seismic and volcanic scores each ≥ 60.
**Hunting access does not open the `personal_defense` gate.** The gate exists for
`firearm_legality` and `distance_to_active_conflict_km` under REVIEW §1 Q5, and hunting is a
lifestyle preference rather than a defence posture; it belongs to the Fit quiz (REVIEW D4), not
to the FCS. `temperate_steady` therefore enables no gates.
---
## 11. Sensitivity and rank stability
REVIEW §1 Q3 closes the weights debate empirically rather than by argument, so the
weights must be shown to be non-load-bearing.
> **The full analysis is published in `docs/SENSITIVITY_REPORT.md`** (re-issued 2026-09-01
> for `0.2.0`, re-measured per QA's rule — never copied; first edition 2026-08-30 at
> `0.1.1`: weight perturbation, indicator dropout, veto sensitivity, confidence×volatility
> over all 7,370 units, plus the A4 calibration at 0.2.0). Headline: 95.3% of units move
> ≤5% of their cohort under any single ±20% weight perturbation (was 95.0%), and no single
> dataset is load-bearing (leave-one-out ρ ≥ 0.9612; the floor is the C1-concentrated
> `conflict_events_per_100k`, the other fourteen ≥ 0.992). The model's largest instability
> source is the veto-cap plateau edge, not the weights — in both editions.
> Published for readers on `/methodology` ("How much these choices matter");
> reproduce it with `uv run marts sensitivity --full` (~350 s).
Procedure (`sensitivity.py`; run over the global cohort at the frozen weights and published in
`docs/QA_REPORT.md` §11 as of 2026-08-30 — `uv run marts sensitivity` writes
`data/marts/sensitivity.parquet`, `sensitivity_summary.json` and
`apps/web/src/data/sensitivity.json`):
1. Score every unit at the baseline persona vector; rank within level (`1` = best,
ties by the `"min"` method).
2. For each of `n_draws` (default **500**) draws: multiply each pillar weight by an
independent `Uniform(0.8, 1.2)` factor (**±20 %**), renormalize to 100, re-score, re-rank.
3. Per unit report: `baseline_rank`, `rank_p05`, `rank_p50`, `rank_p95`,
`rank_swing = rank_p95 − rank_p05`, `max_abs_delta`, `mean_abs_delta`.
4. **`volatile = rank_swing > 20`** — the unit's position is an artifact of the weights
and must be shown with a stability chip rather than as a rank.
**Determinism:** the RNG is `numpy.random.default_rng(seed)`, `seed` defaults to
`20260828`, and **the seed is written into the report output**. Same inputs + same seed
= byte-identical report, always.
### 11.1 Run diagnostics are part of the methodology
`futurecast_scoring.diagnostics` computes the §6.2 per-level veto rate against thresholds held
in the methodology config (`diagnostics.veto_rate_smell`, `diagnostics.veto_rate_tripwire`).
It changes no score. It is here, rather than in a test, because it is **the model's own published
statement of when to distrust its output**, and a threshold that lives only in a test is a
threshold nobody applies: it cannot be quoted on `/methodology`, cannot be printed by the marts
run, and drifts the first time someone edits the test to make it pass.
### 11.2 The rank-shift chip (REVIEW B7)
> How much does the horizon change this place's standing?
`futurecast_scoring.rankshift` compares a unit's rank **within its own level** at
`rank_shift.from` (`now`/`obs`) and `rank_shift.to` (`2050`/`ssp245`) — the two views REVIEW §1
Q8 puts on the free tier — and emits the sidecar mart `data/marts/rank_shift.parquet`. It is a
sidecar rather than columns on `composite_scores` because the chip is a statement about a *pair*
of horizon rows and `composite_scores` is keyed on one of them.
Sign conventions, fixed here so the UI cannot reinvent them:
- `rank_delta = rank_from − rank_to`. **Positive means the unit moved up** (a better rank is a
smaller number), so the sign matches what a reader expects from a chip that says "+9".
- `fcs_delta = fcs_to − fcs_from`. A unit can *lose score and gain rank* — everyone around it
fell further — so both ship and neither is derived from the other.
- `percentile_delta` is the same movement as share-of-cohort-overtaken, and **is the figure to
render whenever levels are mixed**: nine places among 250 countries and nine among 3,600
regions are not the same event.
**Two honesty constraints are in the data, not left to the UI.**
1. **`n_in_level` always ships**, because a rank delta without its cohort size is not a fact.
2. **`material`** — a shift smaller than the unit's own ±20 % weight swing (§11) is weight noise,
not a horizon finding. The chip runs the sensitivity ensemble at *both* ends, takes the worse
of the two swings as the `noise_floor`, and sets `material = |rank_delta| > noise_floor`.
Where no ensemble judged it, `material` is **`<NA>`, never `false`**. A non-material shift
must render as "within the noise", never as "climbs 6 places".
Measured on the Phase-0 320 at `now → 2050/ssp245`: 150 units improve, 91 decline, median
|rank delta| 1, and **31 of 320 clear their own noise floor.** That number is the point of
the chip: it answers "climate matters, but how much *here*" with a figure that is usually
"less than you would guess from the arrows".
---
## 12. Versioning (semver)
`methodology_version` is required input to every scoring run and is stamped on every
output row. Published versions are immutable: a change re-scores everything and writes
new rows (CLAUDE.md §3.2).
| Change | Bump | Gate |
|---|---|---|
| Adding, removing, or reassigning an **indicator**; changing a **formula** or aggregation rule; changing normalization | **major** | **[HUMAN]** |
| Changing **pillar weights**, `weight_in_pillar`, veto threshold/cap, clamp bounds, winsor limits, persona vectors | **minor** | **[HUMAN]** |
| Refreshing source **data** under an unchanged spec; adding units; fixing a bug that does not change any published number | **patch** | agent |
| Freezing a draft (`0.1.0-draft` → `0.1.0`) | release | **[HUMAN]** — *taken 2026-08-30, REVIEW D10* |
| Demoting a restricted indicator and re-weighting its family (`0.1.0` → `0.1.1`, §16) | **minor** | **[HUMAN]** — *taken 2026-08-30, REVIEW D12* |
**`0.1.0` and `0.1.1` are both frozen.** Nothing in this document may be edited to change a
number published under either. The next weight, threshold, formula or indicator-set change
opens a new version, re-freezes the reference library into its own directory, and re-scores.
Corrections to *prose* — a clearer explanation, a recorded fact, a cross-reference — are
allowed and expected; corrections to *rules* are not.
**On `0.1.1`'s number, stated so nobody has to reconstruct it.** The demotions of §16 are a
weight change, which the table above calls a **minor** bump; the string REVIEW **D12** minted
is `0.1.1`. The gate that governs a weight change is a **[HUMAN]** ballot and D12 is that
ballot, so the discipline that matters was kept: a new immutable version string, a new
reference-library directory, a new tile archive, a full re-score, and `0.1.0` left untouched on
disk. The digit is Eric's; the isolation is the rule.
A reference distribution (§4.1) is frozen per **minor** version: any bump that changes
weights or the indicator set rebuilds and re-freezes them. **A re-freeze against a different
cohort is itself a version event**, even under an unchanged version string, because a percentile
rank means "against these units": a library frozen over 320 units and one frozen over 6,400 are
different methodology state. The reference-library manifest therefore records the cohort it was
built from (`n_units`, `units_by_level`) so that claim is checkable rather than asserted.
### 12.1 Cadence (REVIEW addendum 2026-08-29-b, D6)
| cadence | what happens | version effect |
|---|---|---|
| **monthly** | data refresh under an unchanged spec | **patch** — agent |
| **quarterly** | score republication: new version rows, movers and alerts fire | **patch** — agent |
| **annual** | methodology review: weights, indicator set, the v0.2 items named throughout this document | **major/minor** — **[HUMAN]** |
Political-era drift — a Trump-vs-Obama United States — enters through the annual WGI, V-Dem and
Freedom House updates and compounds across versions. It is visible in `score_changes` and is
**never a daily jitter**. News and event streams (GDELT, ACLED when licensed) are a display-only
"pulse" layer on the card in Phase 2; they feed no score until validated against the annual data.
---
## 13. Sanity anchors (acceptance)
These are the acceptance test for the whole pipeline, encoded in
`packages/scoring/tests/test_anchors.py` and run by QA (CLAUDE.md §4). If a data change
breaks one, the data or the methodology is wrong — not the anchor.
1. Norway, Switzerland and New Zealand occupy the top of the Phase-0 ten countries by FCS.
2. Yemen and Sudan occupy the bottom of the Phase-0 ten.
3. Phoenix's Water pillar score < Seattle's.
4. Miami's `coastal_flood_slr_exposure` normalized score is the worst of the Phase-0 US cities.
*(`0.3.0`: worst **or tied-worst** — under anchoring eight US cities saturate at LECZ 1.000 and
all score 0, so the anchor asserts "no US city is worse", which is what it always meant.)*
5. Colorado's FCS > Florida's FCS on the **2050 climate horizon** (SSP2-4.5).
6. Lisbon FCS > Cairo FCS, and Cairo Affordability > Lisbon Affordability (brief §3.3).
7. Massachusetts `life_expectancy_at_birth` > Mississippi's; Colima `homicides_per_100k` >
Yucatán's; Louisiana's > New Hampshire's. (Sub-national variance exists where measured —
the anti-photocopy anchors, promoted from the 0.2 candidate pipelines.)
8. Oceanside and Pittsburgh differ on `pm25_surface_annual`; the five polluted anchor
cities (Delhi, Dhaka, Lahore, Cairo, Beijing) each read > 3× the cleanest of the five
clean anchor cities (Zürich, Oslo, Helsinki, Wellington, Auckland).
9. **Uniform warming moves the climate pillar** (`0.3.0`, ballot E3): the mean climate-pillar
change now → 2050 (SSP2-4.5) is below −1.0 at every level, and the warmest-month heat score
falls for at least 90 % of units. (Measured at the draft: −2.1 L0 / −2.4 L1 / −2.6 CITY;
95.7 % of units.) This is the anchor that would have caught `0.2.0`'s flat 2050.
10. **One scale** (`0.3.0`, ballot A2): a 35 °C warmest month scores 50 at L0, L1 and CITY,
at `now` and in 2050, under both scenarios.
**Dispositions taken at the Phase-0 [HUMAN] gate (Eric, 2026-08-29 —
`docs/PHASE_0_DONE.md`).** Four anchors pass outright. The other two were accepted as
**diagnosed failures**, which is not the same as accepted as passes: each is a `strict=True`
xfail, so the suite fails the day new data makes it pass, and the agent that sees it must read
the diagnosis before recording a win.
| anchor | disposition | why, and what would close it |
|---|---|---|
| 1 · NO/CH/NZ top of the ten | **passes** — NOR 77.5, CHE 74.9, NZL 69.9 | |
| 2 · YE/SD bottom of the ten | **strict xfail** — Yemen is last, **Sudan is 8th** | Not a weight problem. Sudan scores 86.4 on food because the pillar measured land capacity only (§3.5.3), and its homicide value is from **2008**, before the civil war (§3.5.4). Closes on Phase-1 *coverage*: `prevalence_of_undernourishment` (§3.6.3) and `conflict_event_density_50km` (§3.6.2). It was deliberately **not** closed by reweighting. |
| 3 · Phoenix water < Seattle water | **passes** — 15.4 vs 69.8 | |
| 4 · Miami coastal-flood worst US city | **passes** | |
| 5 · Colorado FCS > Florida FCS @2050 | **strict xfail** — accepted as *structurally inexpressible at this resolution* | Colorado and Florida are both US admin-1 units and inherit identical non-climate pillars from `L0:USA`; Miami's LECZ is saturated at 1.000 and cannot worsen under percentile-rank normalization; global-mean SLR moves 43 of 320 units by ~0.005. Removing the veto does not fix it (uncapped CO 49.96 < FL 51.01). Closes on **regional** SLR and state-level pillar data. |
| 5R · Florida's coastal exposure worsens by 2050 | **passes** — the mechanism-true reframe, asserted on the **raw** indicator value, because a saturated percentile cannot move but the physical quantity must | |
| 6 · Lisbon FCS > Cairo; Cairo affordability > Lisbon | **passes** — 47.2 vs 19.8; affordability 95.0 vs 70.0 | The affordability half rests on **one** inherited indicator at confidence 0.06, so it ships with a companion assertion that the Access axis reports its own thinness. |
Anchor 6's affordability half was recorded as *not evaluable* at Phase 0 and became evaluable
when `price_level_index_gdp` landed; the test that asserted the gap was written to fail the day
data arrived, and it did.
**Anchors are not a scoreboard.** If a data change breaks one, the data or the methodology is
wrong — not the anchor (CLAUDE.md §4). Anchor 5 is the case where the *anchor* went to
**[HUMAN]** instead, and was kept rather than deleted precisely so the claim it makes stays on
the record until the model can express it.
---
## 14. What this methodology does not claim
- It is **not** a forecast of 2050 conditions for anything but climate and water, and
even there it is a scenario, not a prediction (§8).
- It is **not** advice — informational only, not financial, legal, or relocation advice
(CLAUDE.md §3.8).
- It makes **no property-level claim**. The finest unit in v0.1 is a city or admin-1
region; a score does not describe a parcel.
- It is **relative, not absolute** (§4.1 known limitation).
- Its heat indicators describe **sustained seasonal heat, not peak lethality**, at v0.1
(§3.3). Threshold-day counts return in Phase 1.
- Ranks near each other are **not** meaningfully different — see the stability chip (§11), and
a horizon rank shift inside that band is **not a finding** (§11.2).
- It makes **no cultural, religious, ethnic or political judgement about the people who live
somewhere**. Those facts appear on the card as context for the reader's own decision and
carry zero weight anywhere in either axis (REVIEW D1). Nothing in the score is a statement
about who belongs where.
- Its climate figures are **scenario-labelled planning ranges in actuarial language** — the way
a lender or insurer talks about risk — not advocacy, and never a single deterministic future
(REVIEW D7).
---
## 15. The freeze record — what was true on 2026-08-30
**This section was a checklist until the freeze. It is now a record, and it is not edited to
look better later.** `0.1.0` was frozen by REVIEW **D10** (Eric's ballot, 2026-08-30), which
also carried **D9** (Blue Zones dropped) from the manual-table review the same day. Every
**[HUMAN]** item below is now decided; each says what was decided and by which ballot line.
An item that was decided *against* closing — accepted as a known limitation rather than
fixed — says so, because "decided" and "resolved" are not the same word.
### 15.1 Engineering — closed
- [x] **Global reference distributions rebuilt and re-frozen** (2026-08-30, wave C): 222
distributions over the 7,370-unit cohort. Re-frozen again at the freeze into
`data/staged/reference_library/0.1.0/` — same 222 keys, same cohort (250 L0 / 3,236 L1 /
3,884 CITY, 591,547 rows), new digest because the version string is stamped in the file.
**Both shas are recorded in `docs/CHANGELOG.md`** and the `0.1.0-draft` library was left
untouched on disk (CLAUDE.md §3.2).
- [x] **Published marts re-scored** (2026-08-30, freeze): every mart, the tiles, the web data
and the rank-shift chip rebuilt at `0.1.0` against the `0.1.0` library, with
`test_marts_on_disk_match_a_fresh_run` green and two runs byte-identical. The marts on
disk before this run were computed under the pre-D9 health vector and did not match a
fresh run — the D9 removal changed four `weight_in_pillar` values and nothing had
re-scored since.
- [x] **`source_spec_mismatch` reads 0** (2026-08-30, wave C).
- [x] **Anchors re-run against the global cohort at the frozen weights**: all 20 pass. The
`hale` demotion moved every health pillar score; no anchor changed verdict. Numbers in
`docs/QA_REPORT.md`.
- [x] **Veto rate re-measured per level** and reported in `docs/QA_REPORT.md` with its
coverage diagnosis.
- [x] **`sensitivity.py` report published** with its seed (§11) — `uv run marts sensitivity`,
seed `20260828`, 500 draws, ±20 %, over the global cohort at the frozen weights.
Per-level volatile shares in `docs/QA_REPORT.md` §11; per-unit flags exported to
`apps/web/src/data/sensitivity.json`. This was the last engineering item outstanding.
- [x] **`docs/CHANGELOG.md` carries the complete freeze diff** — every draft amendment since
2026-08-28, plus the `[0.1.0]` release entry with its known-limitations list.
- [ ] **Every Phase-1 indicator staged** (§3.6) — **not closed, and frozen anyway.** Eight ids
carry weight with no staged rows at all: `rent_yield`, `visa_pathways`,
`passport_mobility_henley`, `foreign_ownership_rules`,
`desalination_capacity_per_capita`, and the rest of the Access set in §3.7. They are
charged to coverage, which is exactly why the Access axis reports the confidence it
does. Freezing with them empty was ratified by D10-A.
- [ ] **`mean_years_schooling`'s carrier registered** — **not closed.** The World Bank
republication sits in EdStats, which `wdi_macro` does not register. Unchanged at the
freeze; the id is still pointed at a source that does not carry it, which is the one
traceability wart in the frozen set and is logged in `docs/DATA_ISSUES.md`.
- [ ] **No scored indicator carries `tier: free_only`** (§3.1 rule 7) — **NOT ACHIEVED.**
D10-B closed `hale`, which was the exception the checklist named. It was not the only
one. `tax_burden` (0.12, `heritage_efi`, 6,951 rows) and `freedom_house_score` (0.04,
`freedom_house`, 7,095 rows) are both scored on `review` carriers. See §3.1 rule 7 for
the full statement and ISSUE-132 for the log. **`0.1.0` ships with two licence
exposures, not zero.**
*(Not an edit to the record: this line stayed NOT ACHIEVED for the whole life of
`0.1.0`. REVIEW **D12** demoted both at **`0.1.1`** the same day — §16.)*
### 15.2 [HUMAN] decisions — all decided, several decided as "accept"
| item | decision | ballot |
|---|---|---|
| **`hale`** | **Demoted to display-only**; its 0.14 → `life_expectancy` (0.25 → 0.39). Ships as an auto-restore candidate if WHO permission lands. | **D10-B** |
| **The food reshape** (§3.6.3) — capacity 0.82 → 0.70, supply security 0.18 → 0.30, `prevalence_of_undernourishment` at 0.14 | **Ratified as documented.** | **D10-A** |
| **Access weights 45/35/20** | **Ratified as documented**, still "for now; tweak with new data" — `legal_access` and `fit` now carry values and the vector was not revisited. | **D10-A** |
| **The licence reconciliation of 2026-08-30** (§3.6.4) — six indicators out, five re-pointed, six pillars re-weighted | **Cost accepted.** Access/fit is two in-house derivations; `legal_access` is 0.76 dependent on PR #1; safety carries no terrorism signal. No licence funded. | **D10-A** |
| **`tax_burden`** | **Not decided by D10** — D10-B addressed `hale` only. Frozen as a scored `review` exposure. **Open, and now the largest single one.** | — |
| **The two IEP indices are gone, not deferred** | **Loss accepted**; no IEP commercial licence authorised. GTI's information is simply absent from safety. | **D10-A** |
| **`english_proficiency` and the thin fit pillar** | **Accepted**: a 20-weight Access pillar built entirely in-house (`winter_severity` 0.55, `timezone_overlap_us_hours` 0.45). No manual table commissioned. | **D10-A** |
| **`vdem_liberal_democracy` and share-alike** (counsel memo Q1) | **Accepted as-is for `0.1.0`**; stability's democratic-institutions signal is WGI plus a `review` Freedom House score. A "no" from counsel restores V-Dem at 0.08 — a v0.2 event. | **D10-A** |
| **`seismic_hazard` ships US-only** (ISSUE-110) | **Accepted.** The frozen safety pillar does not penalise Tokyo, Santiago or Istanbul for seismicity, and only the confidence figure discloses it. This is the limitation most likely to embarrass the product; it is in the release note for that reason. | **D10-A** |
| **GFSI / OECD-RWB display-only** | **Signed off** as written (the GPI half is moot — removed 2026-08-30). | **D10-A**, D3 by analogy |
| **The recency curve and quality multipliers** (§7, §5.2) | **Ratified as documented.** Grace 2 y, decay 8 y, floor 0.5, `inherited_L0` 0.50 — scoring-agent constructions, now [HUMAN]-blessed rather than merely unchallenged. | **D10-A** |
| **The veto is a relative threshold** (§6.2) | **Accepted, framing carried elsewhere.** The maths is untouched for `0.1.0`; the burden moves to card copy and the thin-data chip. **Absolute anchoring is the v0.2 flagship.** | **D10-D** |
| **Published rankings floor** | **New at the freeze**: coverage ≥ 0.40 or no rank, labelled "insufficient data to rank" (§7.1). | **D10-C** |
| **PR #1 (manual tables)** | Merged 2026-08-30 (commit `116078e`), all 56 rows approved by Eric. `legal_access` is no longer frozen empty. | D9 review |
| **`passport_mobility_henley`** | **Not decided.** The id still names a proprietary source and carries 0.12 with no rows. Deprecate-and-replace with `visa_free_destinations` (§3.7) is a v0.2 item. | — |
| **Launch copy and disclaimers** | Still owed. Not a freeze blocker — the methodology freeze and the launch gate are separate **[HUMAN]** gates (CLAUDE.md §2 Phase 1 exit). | — |
| **The product is named Habitance** | `habitance.io` confirmed available 2026-08-30. Paid trademark clearance before any filing. Does not touch this document's numbers; recorded because every surface's footer changes. | **D10-E** |
### 15.3 Explicitly *not* in `0.1.0` — the v0.2 queue as it stood at the freeze
Recorded so the freeze was not held hostage to them, and so the next revision starts from a
list rather than from memory: **absolute anchoring of the physical pillars** (§4.1 — D10-D
names it the v0.2 flagship); `warming_velocity` (REVIEW B1); the 20-year climate window
(§3.4.4); regional rather than global-mean sea level (§3.4.1); splitting `precip_variability`
(§3.4.2); restoring `heat_days_35c` / `wetbulb_days_31c` / `drought_spei` from the daily
archive (§3.3); Gini and the other REVIEW B4 social-capital candidates. Each is a **major**
bump when it lands. Ahead of all of them in priority, because they are corrections rather
than additions: **`tax_burden` and `freedom_house_score`** (§3.1 rule 7) and
**`mean_years_schooling`'s carrier** (§15.1).
The consolidated, ordered version of this queue — across every doc, not only this one — is
`docs/V0_2_QUEUE.md`. This section is the methodology's own slice of it and stays frozen as
written; the queue is the living document.
---
## 16. `0.1.1` — the rule-7 demotions (REVIEW D12, 2026-08-30)
**§15 is `0.1.0`'s record and stays as written.** This section is `0.1.1`'s. It exists because
§15.1 had to record a rule as **NOT ACHIEVED** — *"no scored indicator carries `tier:
free_only`"* — and a rule that the document itself reports broken is not yet a rule.
### 16.1 What D12 decided
ISSUE-132 put three options to **[HUMAN]**: demote both exposures, wait for the outreach, or
accept and disclose. Eric took **demote both**, and kept the outreach alive as the condition
that brings each indicator *back* rather than as a reason to hold the demotion:
| indicator | axis · pillar | weight at `0.1.0` | carrier | verdict | staged rows (kept) |
|---|---|---|---|---|---|
| `tax_burden` | Access · `legal_access` | 0.12 | `heritage_efi` | `review` — no licence instrument published anywhere on the Index site | 6,951 |
| `freedom_house_score` | FCS · `stability` | 0.04 | `freedom_house` | `review` — commercial use requires prior formal permission | 7,095 |
Both keep their id, their staged values and their `tier: free_only` label, and both carry
`weight_in_pillar: 0` with `display_only: true`. **Neither is a statement about the
indicator.** `tax_burden`'s inversion is sound and `freedom_house_score`'s publication-lag
argument (§3.6.2) is unrefuted. Free-tier *display* is what a `review` carrier tolerates, so
both still render on the card, marked.
**Why "wait" was not enough.** `heritage_efi` and `freedom_house` are not *refused*; they are
**unread**, which under rule 7 is the same thing, because the rule turns on what is *cleared*
rather than on what is likely. Waiting is a fine plan for a permission request and a bad basis
for a published number.
### 16.2 The redistributions
Both are within-family, proportional to existing weights, and land on exact hundredths by
largest remainder. No indicator outside an affected family moved, and each pillar's
default-active weights still sum to exactly 1.0.
**`legal_access`** — `tax_burden`'s 0.12 across the pillar's remaining *scored* members
(`property_rights_index` was already display-only and takes none):
| indicator | `0.1.0` | exact target | **`0.1.1`** |
|---|---|---|---|
| `foreign_ownership_rules` | 0.41 | 0.4659 | **0.46** |
| `visa_pathways` | 0.35 | 0.3977 | **0.40** |
| `passport_mobility_henley` | 0.12 | 0.1364 | **0.14** |
| `tax_burden` | 0.12 | — | **0.0 — display-only** |
| `property_rights_index` | 0.0 display-only | — | 0.0 |
| **sum** | **1.00** | 1.0000 | **1.00** |
**`stability`** — `freedom_house_score`'s 0.04 across the WGI family, the same four series and
the same pattern V-Dem's 0.08 took in §3.6.4. Proportional and equal-split agree here to the
hundredth:
| indicator | `0.1.0` | **`0.1.1`** |
|---|---|---|
| `wgi_political_stability` | 0.16 | **0.17** |
| `wgi_rule_of_law` | 0.16 | **0.17** |
| `wgi_control_of_corruption` | 0.12 | **0.13** |
| `wgi_government_effectiveness` | 0.12 | **0.13** |
| `freedom_house_score` | 0.04 | **0.0 — display-only** |
| WGI family total | 0.56 | **0.60** |
| **pillar sum** | **1.00** | **1.00** |
### 16.3 What it cost — read this before reading an Access score
**`legal_access` now scores nothing, for every unit on Earth.** This is the largest single
consequence of `0.1.1` and it is worse than ISSUE-132's option (1) predicted. That option
expected the pillar to fall back on "two manual-table indicators plus `passport_mobility_henley`
(which has no rows)". In fact **none of the three has rows**: PR #1's tables are approved
(§15.2) but were never staged into `indicator_values`, and `tax_burden` was the pillar's only
staged member. So:
* `legal_access` score is **null for all 7,370 units** (it was null for 419 at `0.1.0`);
* its confidence reads **0.00** (median 0.06 at `0.1.0`);
* the Access composite renormalizes over `affordability` and `fit` alone (§4.3, §9), so the
35-weight pillar is dropped rather than scored low;
* every unit's Access number moves — **median |Δ| ≈ 7.9 points, max 38.1** — and the quadrant
(§6.3) flips for a large share of units purely because a pillar left the mean.
**The FCS barely moves**, which is the expected shape: `freedom_house_score` held 0.04 of an
18-weight pillar. Across every unit, persona and horizon the FCS moves by at most **1.09
points**, mean **0.11**. No sanity anchor (§13) changed verdict; all 20 pass.
**Do not read this as "Access got worse."** Access at `0.1.0` was a number in which one third
of the design was carried by a single unlicensed tax index. `0.1.1` declines to publish that.
The honest description of the Access axis today is **affordability and fit, with legal access
unmeasured** — and the fix is staging PR #1's tables, not restoring the demotion.
> **Closed the same day, by staging (§3.6.7, ISSUE-134).** This paragraph is left standing
> because it is `0.1.1`'s record at its mint, and the mint is what §12 makes immutable. What
> followed it, hours later and before `0.1.1` was ever published: `foreign_ownership_rules`,
> `visa_pathways` and `desalination_capacity_per_capita` were ingested from the approved manual
> tables as a **data patch** — no weight, formula or id moved — and `legal_access` went from
> null everywhere to scored for **949 units** (8 L0 + 225 L1 + 716 CITY), confidence 0.688 at
> those L0s. The Access composite is three pillars again wherever the tables reach and two
> pillars everywhere else, which is a coverage statement rather than a design one. The
> `0.1.1` reference library was re-frozen additively (222 unchanged distributions + 9 new keys,
> verified key by key) and both digests are in the CHANGELOG.
### 16.4 What brings them back
Both restorations are indicator-set changes ⇒ **major** bump and a **[HUMAN]** gate (§12). The
word "auto" in "auto-restore" names the *trigger*, never the mechanism.
| indicator | condition | effect |
|---|---|---|
| `freedom_house_score` | a **yes** to the permission request sent **2026-08-30** | `freedom_house` → `license_ok_commercial: true`, `tier: pro_ok`; restore 0.04, return the four WGI series to 0.16/0.16/0.12/0.12 |
| `tax_burden` | **counsel Q23** — the silent-publisher factual-data question, whose Q23(f) is the Heritage contact route | `heritage_efi` → `true`; restore 0.12, return `legal_access` to 0.41/0.35/0.12 |
A "no" to either makes the demotion permanent and the indicator a candidate for
deprecate-and-replace, exactly as `passport_mobility_henley` already is (§3.7).
### 16.5 The gap that let it happen, closed
`0.1.0` shipped these two exposures past a checklist written specifically to find them. The
mechanism is worth naming because it is reusable: **the licence gate checked labelling and
never checked weight.** Both indicators were correctly marked `tier: free_only`, so
`tests/test_license_gate.py` was green while the rule it enforces was broken twice, for ~7,000
units each, in a number served identically to free, Pro and API.
`tests/qa_lib.validate_restricted_indicator_weights` is the second half (§3.1 rule 7, machine
form 2). It reads the **methodology YAML** — the normative home of `weight_in_pillar`, not the
`indicators/*.yaml` mirror — and fails any id whose sources include a carrier that is not
`license_ok_commercial: true` **and** `status: verified` while it holds non-zero weight. It is
fixture-tested against a deliberate violation, and run against `0.1.0`'s own methodology it
reports exactly the two rows in §16.1 — which is how we know it bites rather than merely
passes.
### 16.6 What `0.1.1` is *not*
It is a **weight** change and nothing else. No formula, no threshold, no normalization, no new
or removed id, no change to the reference cohort: the `0.1.1` library holds the same 222
distributions over the same 7,370 units as `0.1.0`, and differs only by the version string
stamped inside the file. Everything on §15.3's v0.2 queue is still on it, minus the two lines
this section closes.
## 17. The `0.2.0` freeze record — what was true on 2026-09-01
Frozen at the stage-3 QA gate the same day the ballot opened it (docs/QA_REPORT.md,
2026-09-01 section, verdict READY TO FREEZE). What the gate verified, in the order the
ballot required it:
- **A4 calibration** (`futurecast_scoring.sensitivity.inheritance_calibration`,
deterministic 8-vector grid over the real 596,571-row 0.2.0 scoring input):
`rank_invariant: True` — every candidate vector left every rank identical
(fcs_max_abs_delta 0.0, Spearman 1.0). Adopted multipliers: `measured_TL2` **0.70**,
`inherited_TL2` **0.70**, `inherited_L0` **0.50** — the strictest inherited_L0 value that
keeps the rankable share above 95% (0.956; 0.40 would have dropped it to 0.842).
`inherited_L1` is currently inert (no staged row carries it).
- **Suites**: packages/scoring 282 · pipelines 141 · pipelines/marts 53 (+1 skip pending
the stage-4 export) · root tests 91 (the one open failure is the §3.10 manual-table gate
correctly refusing the still-unapproved sector-hazard draft, which is not part of 0.2.0).
License gate green at v0_2 — every newly weighted id sits on a verified,
commercially-cleared carrier.
- **Carrier exclusivity (rule 8)**: 0 violations, 0 unattributed rows across 441,682
scored rows.
- **Movement, audited independently**: 1,998 units moved ≥1 FCS point, 106 ≥3, 26 ≥5
(median |Δ| 0.55). All 26 large movers are veto-cap boundary crossings and each traces
to a real sub-20 pillar score appearing or disappearing; the largest uncapped move in
the corpus is 6.47. The Gulf/Sahel health movement matches the §disclosure figures in
`docs/CHANGELOG.md` exactly. Iceland overtakes Norway for the L0 #1 position.
- **Sensitivity** re-measured on 0.2.0: the weights-not-load-bearing result holds and
strengthens (95.3% of units move ≤5% of cohort under any ±20% perturbation, was 95.0%).
The LOO-dropout floor moved to 0.9612 (`conflict_events_per_100k`, a direct arithmetic
consequence of C1's 0.39 concentration) — the published sensitivity report must be
re-measured at its 0.2.0 reissue, never copied forward.
- **The honest cost, stated here as §3.6 demands**: under rule 8, L0 units in TL2-covered
countries lose their national health/safety carriers to the regional ones — 28 countries'
health confidence falls from 1.00 to 0.358 (USA included; safety similarly for 29). The
information got better and the chip fell: the chip measures resolution-weighted coverage,
and a pop-weighted TL2→L0 aggregate under the same source is 0.2.1-ballot business
(deliberately not patched into this freeze).
- **0.1.1 immutability**: sha256 of all six flat mart files, the merged frame, and the
0.1.1 reference library byte-identical before and after the 0.2.0 build.
Known-and-accepted at freeze: the movement mart's `driving_pillar` attributes the
*uncapped* move — every consumer (the /changes ledger, re-narration, movers lists) must
lead with the veto flag when cap state changed, never the driver column.
**Post-freeze data event, recorded (added 2026-09-09).** On 2026-09-02 08:25 the 0.2.0
composite and pillar marts were re-scored once, against the **unchanged** frozen 0.2.0
reference library, after ISSUE-140 retired two geometry artefacts (a 0.0003 km² "Cunene"
sliver and a split Mazandaran) from the unit universe: 7,370 → 7,368 units, 0 non-Mazandaran
values moved, no version event (docs/DATA_ISSUES.md, ISSUE-140). No other file in the
partition has been touched since 2026-08-31. The partition's ten files and the pinned scoring
input are now hashed in `docs/manifests/marts_0.2.0.sha256.json` and verified by
`tests/test_marts_manifest.py` whenever the data is present; "unchanged" is a re-checkable
claim from here on, not an mtime.
## 18. The `0.3.0` freeze record — what was true on 2026-09-10
Frozen 2026-09-09 (`v0_3.yaml`, `frozen_on: 2026-09-09`; ballot A1–E3 and freeze lines F1–F5
signed) and scored 2026-09-10 (docs/QA_REPORT.md, 2026-09-10 section). What the post-freeze
QA pass verified, in the order the freeze required it:
- **The reference library** (`data/staged/reference_library/0.3.0/`, frozen
2026-09-10T14:51:43Z, sha256 `59ecc404…4cbd98`, 1,438,412 bytes): **203 keys over 60
indicators** for a cohort of 7,368 units (250 L0 · 3,234 L1 · 3,884 CITY) and 723,151
scored rows, winsor [2, 98]. Against 0.2.0's 235 keys, 72 were retired — the eight anchored
ids that still had percentile keys at 0.2.0 plus the deprecated `riverine_flood_exposure` —
and 40 were added (the ICP cost ids at L0 only, the two crop-yield variants at 2050 × two
SSPs, `dependency_ratio` at 2050/`wpp_medium`, `road_density`, and the display-only
`landslide_exposure`, `local_warming_delta_c`, `old_age_dependency_ratio`). A curve is not a
distribution: the eleven §4.4 indicators carry no key.
- **Sanity anchors on the real marts (§13), `now`/`obs`, default**: Norway 81.1 > Switzerland
79.3 > New Zealand 77.6 at the top of the Phase-0 ten; Sudan 21.5 and Yemen 18.9 at the
bottom; Phoenix water 43.3 < Seattle 80.1; Miami's `coastal_flood_slr_exposure` 1.0 is the
US maximum (tied with seven others: Cape Coral, Fort Lauderdale, Hialeah, Pembroke Pines,
New Orleans, Chesapeake, Virginia Beach); Colorado 65.1 > Florida-US 61.6 now and 61.3 > 58.9
at 2050/SSP2-4.5 (Florida-UY, 71.2, is a different unit). Lisbon 68.1 > Cairo 33.6 (Cairo
carries a food veto). Iceland 82.5 is the top **rankable** L0; Bonaire-Sint Eustatius-Saba
sits above it at 83.1 with coverage 0.31, below the §7.1 floor, exactly as it did at 0.2.0.
- **Movement 0.2.0 → 0.3.0, computed from the two composite marts** (7,368 units scored on both
sides): mean **+4.6 CITY / +5.0 L0 / +5.4 L1**, 90.2 % of units rise, 3,431 move ≥ 5 points,
417 ≥ 10, largest rise +17.8 (Umm al-Quwain), largest fall −8.3 (Penama, Vanuatu). Two
pillars carry it — **climate median 51.7 → 71.9, water 50.8 → 61.8** (ballot A1–A3) — and two
more moved for reasons the ballot named: health 55.6 → 56.3 (the `dependency_ratio` carrier
hand-over `wdi_macro` → `un_wpp`, 7,276 values) and infrastructure 50.2 → 49.9 (`road_density`
from GRIP4 joins the pillar at 0.06). Community, food, safety and stability are unchanged to
±0.05. Climate-pillar vetoes fell 307 → 4 (A4's promise, measured), water 466 → 290.
- **The veto cap, explained rather than tuned.** `veto_applied` fell 3,009 → 2,857, but the
number of units *published at exactly 40.00* rose 756 → 1,351: 652 units whose raw composite
sat at 27.6–48.3 (median 36.5) rose to 40.0–51.7 (median 42.5) while a ranked-pillar flag —
community 417, health 136, safety 129 — stood still. The flagged pillars did not move; the
raw score crossed 40 under them. 534 of the 652 have an identical flag set to 0.2.0, 102 a
strict subset (a climate or water flag released, another remaining), 16 gained a health or
infrastructure flag. 57 units left the cap because their only flag (climate 23, water 33,
infrastructure 1) cleared.
- **The largest movers, traced.** The UAE's cities and regions +17.0 to +17.8 and the Mekong
delta provinces +16.9 to +17.7: their climate pillar goes from the bottom of a percentile
cohort (7–13) to what the frozen curves say the same values are worth (37–46), F1's 5 cm
riverine floor included; Viet Nam (L0) +16.85 with its cap released. The two falls: Penama
and Sanma (Vanuatu) −8.3 / −7.9, infrastructure 21.8 → 19.7 and 20.4 → 18.4 when `road_density`
arrived — a sub-20 crossing, so the cap. Rotterdam −2.2 and The Hague −2.1 are the
unprotected-hazard riverine curve (ISSUE-157), disclosed in §4.4's basis.
- **Sensitivity re-measured on 0.3.0** (500 draws, ±20 %, seed 20260828): the volatile share
fell 0.924 → **0.765** (5,636 of 7,368; L0 9.6 %, L1 77.3 %, CITY 80.1 %), median rank swing
90.1 → 83.0, median swing as a share of level 2.0 % L0 / 2.3 % L1 / 2.7 % CITY. The weights
are less load-bearing than at 0.2.0, which is what replacing a rank with a curve should do.
The rank-shift mart (now → 2050/SSP2-4.5) marks 2,130 of 7,368 moves material; 3,673
decline, 3,634 improve, 61 hold.
- **Access at 0.3.0**: `legal_coverage` unchanged (0 · 0.43 · 0.688; 949 covered units, eight
countries). The axis ceiling rose 0.556 → 0.637 because the ICP 2021 price level indices were
staged (`wb_icp_2021`, at L0 only, `measured_at_unit`), not because legal data arrived; the
arithmetic ceiling given what is staged is 0.753 before recency. 2,843 units carry a
Move / Buy or Trap verdict and 2,765 of them (97.3 %) with a legal-access confidence of 0.
**Known at freeze, not patched.** (1) The horizon-frames module drops a modelled 2050 row
whose unit has no 2035 anchor (`horizon_frames._blend` at weight 1.0): two cities (Oceanside,
Escondido) show a 2050 water pillar of 79.3 in the frame against 59.0 in the mart, 3.6 FCS
points — the mart is right, the frame is wrong, and the §8 guarantee that the 2050 frame *is*
the published 2050 number is violated at those two units until it is fixed. (2) The two ICP
cost ids reach 172–173 countries and no sub-national unit, unlike `price_level_index_gdp`,
which the national fallback inherits everywhere — so a country's affordability confidence can
read 0.772 while its own cities read at most 0.27. A flag question was checked and closed:
`measured_at_unit` is honest on those rows because nothing inherits them.
---
## 19. Display rules — what a surface derives from the published numbers, changing none of them
*Added 2026-09-10 by the place-page makeover (`docs/plans/2026-09-10-place-page-makeover.md`
§2.1, §6 lane C). **This section carries no version bump.** It defines no score, weight,
threshold, normalization or indicator; it reads only numbers already published under `0.3.0`
and says which of them a surface shows first and in what words. §12 classes that as prose and
display — the same class as the rank-shift chip's sign conventions (§11.2) and the horizon
labels (§8). Changing a rule here is a display change recorded in `docs/CHANGELOG.md` under
"display only"; it never re-scores.*
A display rule exists because an editorial choice a reader cannot reconstruct looks like an
opinion. Every choice below is deterministic, reads named export keys, and can be re-derived
from `apps/web/src/data/units.json` by anyone who disagrees with it.
### 19.1 The lead figure at 2050 — which pillar the place page shows first
The place page opens on three figures: the score now, the one pillar that moves most by 2050,
and the legal answer. The second is a selection among published numbers, and this is the rule.
**Candidates.** The pillars whose spec declares a `projection` block with a 2050 horizon
(§8; `apps/web/src/lib/schema.ts` `PILLARS[].projection`) — read from the schema, never a
hand list. At `0.3.0` that is four: climate and water (`family: climate`, scenario SSP2-4.5,
keys `p1_50_245` / `p2_50_245`), food (`family: climate`, SSP2-4.5, `p8_50_245`) and health
(`family: demographic`, UN medium, `p4_50_wpp`). A pillar missing either its `now` byte or
its 2050 byte on the unit is not a candidate.
**Selection.**
```
delta_p = score_p(2050, family scenario) − score_p(now) # integer bytes in the export
lead = argmax_p |delta_p|
tie-break = the larger default weight (§2: climate 25 > water 18 > health 10 > food 5)
hold = max_p |delta_p| < 2
```
The tie-break cannot itself tie because the four default weights are distinct. The scenario
is the family's own (§8): SSP2-4.5 for the climate family, the UN medium variant for the
demographic family — never a request-side SSP relabelled onto health.
**What renders.** The lead pillar's two scores, `now → 2050`, on the ramp (both are scores;
the arrow is ink), and a caption that carries, in this order: the pillar and the horizon
(`{Pillar} by 2050`); the **physical row** where one exists; the composite at 2050 with **its
own** confidence byte (`computeFcs(unit, "2050", "ssp245").confidence`, never the `now` byte
relabelled — **and at 2026-09-10 that is exactly what the client would print**: the web
export carries per-pillar confidence bytes `c1..c8` for `now` only (`schema.ts`), so the
client-side 2050 byte equals the `now` byte — 74 for Portugal, against the 0.69 the
narration mart computed for the same horizon. Until a 2050 byte reaches `units.json`, a
surface prints the `now` byte labelled as today's coverage and never as the horizon's; the
gap is a lane-P prerequisite, not a display choice); the rank-shift sentence; and the family label from `pillarHorizonLabel` (§8 —
`2050 climate horizon (SSP2-4.5)` or `2050 demographic horizon (UN medium)`).
- **The physical row** is the ballot E1 requirement of §8 applied to one line. It comes from
`horizon-change/<slug>.json`, restricted to the rows that belong to the lead pillar
(water → `projected_water_stress`, paired with `baseline_water_stress` as `now`; climate →
`tasmax_warmest_month`, `heat_degree_months_30c`, `precip_variability`,
`coastal_flood_slr_exposure`), choosing the row with the largest
`|delta − cohort_median_delta|` from `horizon-change.meta.json` for the unit's level. It
is rendered as the quantity in its own unit with the cohort median beside it
(`baseline stress 0.70 → 0.98; the median country +0.03`). **Food and health have no row
in that shard**, so when either leads the caption carries the pillar delta and its family
label only. No physical quantity is invented for them.
- **The rank-shift sentence** is §11.2's mart in words: `{rank_from} → {rank_to} of
{n_in_level}`, then `, outside the model's own weight noise` when `material` is true,
`, within the model's own weight noise` when false, and no clause when `material` is
`<NA>`. The bare rank delta is never printed (the `data/README.md` rule): a delta without
its cohort and its noise floor is not a fact.
- **The hold state.** When no candidate moves 2 points or more, the row shows the `now`
score in ink and reads `Holds by 2050 — no projected pillar moves more than 1 point`,
followed by the composite at 2050 with its confidence and the horizon label. The printed
figure is `threshold − 1` and, because the export's pillar values are integer bytes
(checked: 0 non-integer values across the eight keys on 7,368 units), it is true by
construction.
- **No candidate at all** (a unit with no projected pillar) → the row carries the pillar-only
caption. 0 units are in that state at `0.3.0`.
**Versions are named where they differ.** `horizon-change/<slug>.json` is stamped
`methodology_version: 0.2.0` today while `units.json` and `rank_shift.json` are `0.3.0`; the
footnote under the row prints the shard's own version beside the page's, never one for both.
**Measured on `0.3.0`, 2026-09-10, all 7,368 units** (the consequence of the rule, disclosed
rather than tuned away):
| lead pillar | L0 | L1 | CITY | all |
|---|---|---|---|---|
| food | 104 | 1,671 | 2,392 | **4,167** |
| water | 73 | 942 | 1,075 | **2,090** |
| health | 46 | 348 | 303 | **697** |
| climate | 13 | 212 | 101 | **326** |
| hold (largest Δ under 2) | 14 | 61 | 13 | **88** |
625 units tie on |Δ| and are settled by weight; the largest tie class is food–water (238),
which resolves to water. **Food leads on 56.6 % of pages** because the `0.3.0` crop-yield
projection (ballot C2) moves the food byte further than the anchored climate and water curves
move theirs for most units; the rule takes the pillar that moves most in the reader's terms and
does not weight the movement, so this is what it shows. The corollary for the physical row: a
climate or water lead carries one on 2,416 units; a food or health lead carries none on
4,864. Both figures are re-derived at every export and are not a property of this text.
**Why `max |Δ|` and not `weight × |Δ|`.** A weighted variant surfaces the 25-weight climate
pillar on nearly every page and turns the row into the composite's own movement restated; the
unweighted variant surfaces the thing that changes most about this place, which is what a
reader who has seen the composite and nothing else does not yet know. The composite's own movement is on
the same line (`Habitance at 2050: {n} /100`), so nothing is hidden by the choice. If
[HUMAN] prefers the weighted form, that is a display change under this section, not a
methodology version.
**What this rule is not.** It is not a score, a ranking or an indicator; it changes no
published number and no confidence; it does not decide which pillars are projected (§8 and
CLAUDE.md §3.5 do); and it is implemented once, as a pure function
(`apps/web/src/lib/leadFigures.ts`, unit-tested) that every surface showing a lead figure
must call rather than re-derive.
**Worked example, Portugal (`L0:PRT`, re-derived at ship).** Climate 74 → 73 (−1), water
63 → 52 (−11), food 49 → 38 (−11), health 82 → 79 (−3). Water and food tie at 11; weight
18 > 5, so water leads: `63 → 52 · Water by 2050 · baseline stress 0.70 → 0.98; the median
country +0.03 · Habitance at 2050: 64 /100, Confidence 69 · 41st → 53rd of 250, outside the
model's own weight noise · 2050 climate horizon (SSP2-4.5)`. The composite figures were
re-derived through `computeFcs` on 2026-09-10 (now 68, confidence 74; 2050 64); the 69 is the
narration mart's and is the figure the export does not yet carry. The rank-shift mart carries
`rank_from 41, rank_to 53, n_in_level 250, material true`. Colorado (`L1`) reads 751 → 964 of
3,234, `material false`, so its sentence says *within*.