Briefing hub for the UN System Data Commons hackathon project
Every other analysis here runs city → UN: take an SDG indicator, find the municipal dataset that matches it. That can only discover what the framework already asks about. This runs it backwards — take every dataset a city publishes, find its nearest SDG indicator, and look at what is left over.
5,956 datasets from 11 city portals (language non-en), matched against all 519 enumerated SDG indicators — not the 442 with usable data, because a framework gap is a question about vocabulary rather than coverage.
What a result here means. A dataset far from every indicator means no SDG indicator’s text is near this dataset’s text. That is evidence about vocabulary, not proof of a conceptual gap — this project has already learned once that a null from the matcher is not evidence of absence. So the unit of evidence below is how many independent cities a theme appears in. A theme in thirty city catalogs is a category of municipal governance; a theme in one is that city’s filing habit.
Datasets from the hand-verified NYC crosswalk. These are known to correspond to an SDG indicator, so they must land in the high-affinity region. If they fall in the tail, the tail is measuring retrieval failure and nothing below is trustworthy.
11 of 11 located · 0 fell in the tail.
| Expected correspondence | City | Dataset | Affinity | Percentile | In tail? |
|---|---|---|---|---|---|
| Population structure (demographics) | Rostock | Bevölkerungsstruktur 2005 | 0.633 | 91 | no |
| Road traffic accidents (3.6.1) | Madrid | Accidentes de tráfico con implicación de b | 0.535 | 90 | no |
| Road traffic accidents and injuries (3.6.1) | Milan | Mobilità: incidenti stradali e persone inf | 0.516 | 90 | no |
| Municipal waste collected (11.6.1) | Milan | Rifiuti raccolti a Milano e nei comuni lim | 0.489 | 88 | no |
| Municipal waste collection (11.6.1) | Belo Horizonte | Coleta de Resíduos | 0.474 | 79 | no |
| Monthly road traffic accidents (3.6.1) | Matera | Bilancio mensile incidenti stradali anno 2 | 0.477 | 72 | no |
| Air quality, hourly (11.6.2) | Madrid | Calidad del aire. Datos horarios desde 200 | 0.462 | 71 | no |
| Water quality (6.3.2) | Buenos Aires | Calidad del Agua | 0.418 | 61 | no |
| Drinking water quality (6.1.1) | Karlsruhe | Jahresmittelwerte zur Trinkwasserqualität | 0.408 | 57 | no |
| Deaths from road traffic accidents (3.6.1) | Fortaleza | Número de Óbitos por Acidentes de Trânsito | 0.333 | 44 | no |
| Road traffic accidents with victims (3.6.1) | Recife | Acidentes de Trânsito com Vítimas 2016 | 0.394 | 32 | no |
The bottom 25% of each catalog by affinity — 1,516 datasets. Below are the title phrases that recur across that tail, counted by how many independent cities use them. These are not inferred categories; they are what the cities themselves called the data.
| Cities | Datasets | Phrase | |—:|—:|—|
No phrase recurs across four cities. Expected for a run pooling five languages — Spanish, Italian, Portuguese, German and Croatian titles share essentially no bigrams, so this view is uninformative here by construction rather than because nothing recurs. The clusters below carry the result.
A second view, kept because it groups datasets that share no vocabulary. k-means returns k clusters whether or not k themes exist, so each carries its coherence — the mean cosine of its members to its own centroid. Below 0.62, a cluster is a partition rather than a theme and is marked diffuse; 12 of 20 are. Read those as noise, not as findings.
| Coherence | Cities | Datasets | Terms | Closest SDG indicator |
|---|---|---|---|---|
| 0.895 | 6 | 241 | elezioni, sezione, risultati, referendum, abrogativo | Countries that have national urban policies |
| 0.861 | 5 | 113 | omi, quotazioni, immobiliari, semestre, locazione | Performance index of data Infrastructure (Pi |
| 0.711 | 3 | 71 | orari, categoria, accessi, area, storica | Countries that have national urban policies |
| 0.698 | 5 | 33 | telelavoro, personale, indeterminato, tempo, anno | Proportion of project objectives of new deve |
| 0.679 | 4 | 51 | pirf, zeis | Degree of implementation of integrated water |
| 0.676 | 2 | 51 | milanesi, accademico, universit, nelle, anno | Open Data Inventory (ODIN) Coverage Index |
| 0.644 | 7 | 39 | covid- | Number of new HIV infections per 1,000 uninf |
| 0.633 | 6 | 69 | ipem | Universal health coverage (UHC) service cove |
| 0.62 (diffuse) | 2 | 26 | geologia, sica, lote | Total greenhouse gas emissions (including la |
| 0.611 (diffuse) | 10 | 69 | (no term covers a fifth of this cluster) | Average proportion of Terrestrial Key Biodiv |
| 0.599 (diffuse) | 7 | 11 | kvaliteti, zraka, zagrebu, podaci, gradu | Carbon dioxide emissions per unit of GDP at |
| 0.594 (diffuse) | 10 | 90 | (no term covers a fifth of this cluster) | Death rate due to road traffic injuries |
| 0.585 (diffuse) | 11 | 120 | (no term covers a fifth of this cluster) | Total public expenditure per capita on cultu |
| 0.557 (diffuse) | 9 | 72 | (no term covers a fifth of this cluster) | Proportion of municipal waste recycled |
| 0.552 (diffuse) | 7 | 70 | licita, fortaleza | Countries that have national urban policies |
| 0.541 (diffuse) | 7 | 69 | zagreba, grada | Proportion of countries with alignment of Na |
| 0.535 (diffuse) | 4 | 35 | (no term covers a fifth of this cluster) | Countries that have national urban policies |
| 0.51 (diffuse) | 7 | 34 | (no term covers a fifth of this cluster) | Countries that have legislative, administrat |
| 0.503 (diffuse) | 9 | 98 | (no term covers a fifth of this cluster) | Countries with procedures in law or policy f |
| 0.469 (diffuse) | 9 | 154 | (no term covers a fifth of this cluster) | Open Data Inventory (ODIN) Coverage Index |
elezioni, sezione, risultati, referendum, abrogativo — 241 datasets across 6 cities (coherence 0.895, mean affinity 0.294)
Most central to the cluster:
One per city, to show the spread:
omi, quotazioni, immobiliari, semestre, locazione — 113 datasets across 5 cities (coherence 0.861, mean affinity 0.32)
Most central to the cluster:
One per city, to show the spread:
orari, categoria, accessi, area, storica — 71 datasets across 3 cities (coherence 0.711, mean affinity 0.306)
Most central to the cluster:
One per city, to show the spread:
telelavoro, personale, indeterminato, tempo, anno — 33 datasets across 5 cities (coherence 0.698, mean affinity 0.316)
Most central to the cluster:
One per city, to show the spread:
pirf, zeis — 51 datasets across 4 cities (coherence 0.679, mean affinity 0.261)
Most central to the cluster:
One per city, to show the spread:
milanesi, accademico, universit, nelle, anno — 51 datasets across 2 cities (coherence 0.676, mean affinity 0.312)
Most central to the cluster:
One per city, to show the spread:
covid- — 39 datasets across 7 cities (coherence 0.644, mean affinity 0.307)
Most central to the cluster:
One per city, to show the spread:
ipem — 69 datasets across 6 cities (coherence 0.633, mean affinity 0.273)
Most central to the cluster:
One per city, to show the spread:
python3 probe/fetch_municipal.py
python3 probe/inverse.py