How district data is reconciled
Census 2011 and NFHS-5 don't use the same district names, and both predate several districts that exist today. This page documents how each source is matched onto the app's current 785-district geography — including where that method is weakest.
The target geography
Every layer resolves to the same 785 districts, sourced from the Local Government Directory boundary releaserather than any one source's own snapshot. That release is itself dated — a small number of districts created since (in Madhya Pradesh, Arunachal Pradesh, Andhra Pradesh, Goa and Ladakh) aren't yet in it, and are mapped to their parent district in the meantime.
Matching by name, not by code
Census and NFHS-5 district names are matched to the current district list directly where the spelling already agrees, which clears the large majority of rows on its own. The rest go through a hand-built alias table — renamed districts (Gurgaon → Gurugram), transliteration differences, and reordered names — rather than a fuzzy best-guess match. A name that still doesn't resolve is left unmatched rather than guessed at.
Districts created after their source was published
Where a district didn't exist when Census 2011 was taken, its historical predecessor is determined geometrically — by intersecting the current district's boundary against 2011's — rather than by hand. Counts (population, households) are then apportioned across the successor districts by land area; rates and survey estimates are inherited unchanged from that predecessor instead, since a rate can't be meaningfully split by area. Every value produced this way is flagged in the data, not shown as an original measurement.
What's actually loaded
Census 2011 contributes population, density, literacy and 15 more demographic fields. NFHS-5 — a survey with roughly 115 published indicators — contributes the 7 with the most complete district coverage; the rest weren't imported rather than backfilled with a weaker estimate. Both are one-time imports, not a live feed, so a change upstream won't appear here until the app is reseeded.
Limitations
- Apportioning a parent district's counts by land area assumes population is spread evenly across it, which it rarely is. Measured against WorldPop's 2020 raster, 33 of the 151 post-2011 districts come out more than 1.5× off and 14 more than 2× — worst where a forested or mountainous carve-out takes a large share of the land and a small share of the people. These apportioned counts are shown on the map layers, flagged as estimated, but are withheld entirely from the individual district pages, which print “not available” instead. Rates are unaffected: inheriting a predecessor's percentage is a defensible proxy where inheriting its population is not.
- A district that inherits its predecessor's rate is identical to it on that indicator, which understates real variation between the two if treated as independent observations.
- NFHS-5 is a sample survey, not a census — district-level figures carry sampling error that isn't currently shown, and are least precise for the smallest districts.
- The highway network's total length figure double-counts stretches built as dual carriageways, since each carriageway is a separate segment in the source data.
The railway zone map is derived from stations
Railway zone boundaries are operational, not geographic, and no zone polygon is published anywhere. The choropleth is therefore built from where each zone's stations are, in five steps, each seeing only the districts the ones above could not answer: a district's own stations (631 districts), a tie-break using the surrounding districts, the nearest zoned stations where track runs through but no station carries a zone (18), a zone's documented territory where there is no railway at all (132), and a small number of hand-checked overrides. Districts in the fourth group are drawn in a muted tone of the zone colour, because they have no railway in them. Every zone present in a district is stored with its station count, so the sidebar can show the whole vote rather than only the winner.
Route kilometres, not track kilometres
Adding up the length of every line in a district more than doubles the real figure, because a four-track corridor is drawn as four parallel lines and every named route that shares it is drawn again on top. So the figure shown is routelength: the geometry is buffered by 20 m and unioned, which dissolves parallel tracks less than 40 m apart into one polygon, and the length is read back off that polygon's perimeter. The width was chosen by measuring against published lengths rather than reasoned from theory — Konkan comes out 1.7% short, the Kashmir line 0.6% long — and widths of 40 m, 60 m and 100 m moved every figure by under 1%. What remains is a few percent of overshoot on long corridors, from branch and junction spurs whose perimeter counts but whose length is not part of the trunk route.
Two sources for the railway, doing different jobs
OpenStreetMap says wherethe railway is — track, stations, named lines, coordinates. Indian Railways' own station master says whoseit is. The second is needed because OSM's operator tag is a volunteer's copy of an administrative fact and goes stale: South Coast Railway was carved out of South Central in 2019, and OSM still tags most of its stations to the old zone. Matching the two by station code corrects the zone on several hundred stations and supplies a division for thousands more. A station's original OSM value is kept alongside, so the two sources never get confused.
A train's route is derived, and checked
No free, current source publishes the track a named service runs on. Open bulk datasets predate the Vande Bharat fleet entirely — the government timetable was last updated in 2018 — and OpenStreetMap carries a route relation for only a minority of services. So every route here is routed over our own track data, and what the source does publish is used to constrain it.
How much it constrains varies by service, so the sidebar names which of three methods drew each line. Strongest is along its named rail lines: some articles list the rail sections a service runs over, and the route is made to pass through the junctions those sections meet at. Next is through its stops, where the calling pattern is published but the track between two consecutive stops is still our inference. Weakest is termini only — nothing published constrains the middle, so the line drawn may not be the one the train runs on.
Every method a service qualifies for is routed, and the published distance picks between them rather than the ranking above: a named-lines route that comes out further from the stated distance than the stops route loses, because agreeing with the source is the evidence and the method is only a description of how we got there. The winner is kept only when its length agrees to within 15%. A route further out than that is stored without geometry rather than drawn along a path the train does not take — which is what happens to a service that deliberately runs the long way round, since a shortest path cannot reproduce one.