The false comfort of a precise number
I built an atlas of ~3,000 battles to finally see military history laid out in space. The hard part wasn't the maps — it was refusing to pretend I knew things I didn't.
By Andrew Pyle
I've watched a lot of war documentaries, and there's a moment in almost every one where I quietly lose the thread. The narrator says the army swung south to flank the enemy at some river I've never heard of, and I nod along, and I have no real idea where any of it is. Not the river, not the flank, not the distances involved. I'm being told a story about space, in a form that shows me none of it.
That gap is what became The War Atlas. Military history is overwhelmingly geographic — terrain, distance, supply lines, who held the high ground — and we mostly consume it as prose, which is the one medium that can't show you space. I wanted to fix that for myself first: to see a whole war at once and then zoom down to a single day's fighting, on a real map, at real coordinates. Strong writing and strong visuals, so the thing finally becomes something you can picture instead of a blur of unfamiliar place names.
01Reader to claimant
From reader to claimant
Here's what I didn't expect. The moment you decide to put history on a map, you stop being a reader and become a claimant. A paragraph can say a battle happened “near” a town and move on. A map pin cannot be “near.” It's at a coordinate, or it isn't anywhere. Every battle suddenly needed a location, a date, a set of numbers — and the second I started demanding those, the sources stopped cooperating.
Because history doesn't come in clean rows. Two respectable sources will give you two casualty figures for the same battle, off by thousands, and both are “right” in the sense that each is citing something real. Ancient battles arrive with numbers that are frank propaganda. Even well-documented modern ones have holes. If you want a tidy database, you have to paper over all of it — pick a number, drop the ambiguity, present a clean surface. That's the standard move. It's also a lie.
02Ranges, not precision
Ranges instead of fake precision
So I built the opposite. On The War Atlas, force strengths and casualties are shown as source-cited ranges with a confidence grade — not single hard figures, because single hard figures are usually fiction. If the evidence supports “somewhere between 23,000 and 28,000,” that's what you see. If a battle has no reliable numbers at all, it shows none; nothing is imputed, nothing is invented to fill the column. Fake precision is the original sin of data, and I wanted a history site that refused to commit it.
Gettysburg is the example I point people to. Union casualties there are recorded almost to the man — 23,049 — because the Union kept meticulous records and they survived. Confederate casualties are a band, roughly 23,000 to 28,000, because their records were thinner and much of what existed was lost in the collapse of the Confederacy. On most sites someone picks a Confederate number and moves on. On mine, you see the gap. And the gap isn't a flaw in the data — it's information. The asymmetry between a precise Union figure and a fuzzy Confederate one is itself a fact about the war and how it ended. Precision and its absence are both telling you something, if you're willing to show both.
03Wrong at scale
Wrongness at scale, including my own
The hard part is that a system built to map everything can also be confidently, authoritatively wrong — and wrongness at scale wears a suit and tie. Early on, the atlas would sometimes show you the wrong battle. Ask for the Battle of Kruty from 1918 and it might hand you a different Battle of Kruty from 2022, because the two shared a name and the pipeline had resolved the reference by guessing from the title instead of following the exact identifier. Nothing looked broken. The page was clean, mapped, cited. It was just about the wrong event. That's the real danger of a polished data product: it makes its errors look like facts.
I try to hold myself to the same standard I hold the data to, which sometimes means catching my own overconfidence. At one point I flagged a defect and reported it as affecting 880 records — then, checking my own detection query, found it had a counting bug of its own, and the real number was 187. I'd done a miniature version of the exact thing the whole project exists to prevent: stated a number with more confidence than it had earned. The fix was easy. The lesson was the point.
04Admit the unknown
Honesty about what you don't know
After enough of this, here's what I actually believe, and it runs a little against the grain: a dataset that openly admits what it doesn't know is more trustworthy than one that looks complete. Completeness is usually a costume. Most “authoritative” data is full of quiet guesses dressed as facts, and almost nobody grades their own confidence, because confidence sells and hedging doesn't. But the hedge is where the honesty lives. Show me your ranges and your blanks, and I'll trust your hard numbers more, not less.
The other thing you learn building this is that the hardest problem in history-as-data isn't the facts — it's identity. There are four Battles of Thermopylae. Three of Mantinea. Wars share names, commanders share names, towns get renamed and re-fought over across centuries. Getting a number right is nearly solved compared to being certain which Kruty, which Thermopylae, which war you're even talking about. Most of the real engineering went there — into being sure of the referent before ever touching the figure.
05What it's for
What it is, and what it's for
I want to be honest about what this is and isn't. It's not primary research. The raw material is largely normalized from Wikidata and Wikipedia — public, open provenance — cross-linked into one relational catalogue and put on a map. A skeptic could call it Wikipedia on a map, and there's a grain of truth in that. But taking thousands of messy, inconsistent, contradictory sources and normalizing them into a single catalogue that grades its own certainty, resolves its own identity collisions, and can be downloaded openly under a real license — that normalizing is the work, and it's harder and more valuable than it sounds. The facts were already public. An honest structure over them wasn't.
People ask what it's for, commercially, and the honest answer is that I didn't build it as a business. I built it because I wanted it to exist and nobody had made the version I wanted to use. I'd like to monetize it eventually — educators, embeds, the open dataset all point somewhere — but that's icing. The value and the interest have to come first; the money, if it comes, comes downstream of having built something genuinely good. Getting that order backwards is how you end up with one more fake-precise history site optimized for ad impressions.
So that's the atlas: a way to finally see a war laid out in space, and a way to trust what you're seeing — including, especially, the places where it says we don't know exactly. Those turned out to be the same project. You can't honestly show someone where a battle happened without also being honest about how sure you are that it happened the way the number says. The map makes it visible; the confidence grade makes it trustworthy. I built it to understand these wars myself — and the understanding and the honesty were never really separable.