September 8, 2026·6 min read

The Soft 404 You Can't See: When HTTP 200 Is Worse Than 404

By Andrew Pyle

A soft 404 is a page that has nothing on it but still returns HTTP 200 instead of 404. Google crawls it, finds a thin or empty shell, and has to guess whether it is a real page or a broken one. That guess is the whole problem, and I shipped one on my own site without noticing.

01

The bug lived in one line of nginx

My /writing section is a React app, prerendered to static HTML per essay. Publish an essay and it gets its own file on disk: `/writing/<slug>/index.html`, full content, right title, ready for a crawler that never runs JavaScript.

The nginx rule that serves it looked like this:

``` try_files $uri $uri/index.html /writing/index.html; ```

Read it as a fallback chain. Try the exact file. Try the file plus `index.html`. If neither exists, fall back to the /writing index page instead of failing. That last clause is the bug. It does not fail. It serves the /writing index shell, a roughly 247-word near-empty page, and it serves it with a 200.

02

I could visit a page that doesn't exist and it would look fine

I checked a slug I made up: `/writing/this-slug-does-not-exist-xyz`. No essay by that name exists and never has. The server returned 200 anyway, with the /writing index page's markup and its title, "Writing, Andrew J. Pyle," instead of anything specific to the slug I typed.

That is the part that makes a soft 404 dangerous: it looks correct in a browser. The page loads, nothing is visibly broken, there is no red error banner. You have to check the status code to know it lied. I only found it because I went looking, not because anything alerted me.

03

The surface is unbounded, and that's the actual danger

An unpublished essay returning the index shell at 200 would already be a minor nuisance. What I found is worse: any string after `/writing/` returns 200. Not just draft slugs waiting to be published. Any slug. Typos, guesses, old links, a scraper's fuzzed URL list, anything.

That means the number of soft-404 URLs on this fallback is not "however many drafts I have." It is unbounded. Every crawlable permutation of `/writing/<anything>` is a distinct URL that Google can discover, request, and receive a 200 for. Each one comes back as the same near-empty, wrong-titled page. To a crawler doing thin-content and near-duplicate detection, that looks like an infinite stack of interchangeable junk pages sharing one domain.

My site is still working back from a Helpful Content demotion. I have written before about pruning the actual thin pages that caused that, and about the discipline of noindexing rather than deleting so a recovery hypothesis stays reversible. An unbounded soft-404 surface undoes that work quietly, from underneath, on a part of the site I wasn't watching, because none of those junk URLs were pages I ever created. They were manufactured by a nginx fallback clause, one request at a time, forever.

04

Why 200 is worse than 404 here

A real 404 is an honest signal. It tells Google: nothing is here, do not index this, drop it from consideration. Google trusts that and moves on. A soft 404 tells the opposite lie: here is a page, please index it. Google has to spend crawl effort figuring out that the page is thin and probably not worth keeping, and until it figures that out, the URL sits in the index as a duplicate of every other soft-404 URL on the same fallback.

The status code is the contract between your server and the crawler. A 404 keeps the contract honest even when the content is missing. A 200 on a missing page breaks the contract while looking, to a human glancing at the rendered page, exactly like nothing happened at all. The lie is the problem, not the missing content. Missing content is normal. A server that claims missing content exists is not.

05

The fix was one line

Every other personal section on the site, /work, /about, /explorations, already used the strict form of the same fallback:

``` try_files $uri $uri/index.html =404; ```

`=404` instead of a soft redirect to an index page. No fallback shell, no guessing. If the file is not there, say so. /writing was the lone exception, the one section still leaking shells, because it had been built slightly earlier and never brought in line with the pattern the rest of the site already followed.

I changed the one line and deployed it. That was the entire fix. No new middleware, no route table, no application code. The config already had a correct pattern sitting three sections over; /writing just wasn't using it.

06

I verified both sides after deploy

I did not take the fix on faith. After deploying I hit a real, published essay slug and confirmed it still returned 200 with the full essay content, headline and all. Then I hit the same made-up slug from before, `/writing/this-slug-does-not-exist-xyz`, and got a real 404 this time. No shell, no 200, no page to index.

That pairing matters more than either check alone. A fix that breaks real essays while closing the soft-404 hole is not a fix, it is a new outage. A fix that leaves the soft 404s and only feels safer is not a fix either. You need both checks, on the actual served response, not on what renders in a browser tab.

07

The broader lesson: check the status code, not the screen

A soft 404 is invisible by design. The page looks fine because it is, technically, a page: it has HTML, it has some words, a browser paints it without complaint. You will never catch one by clicking around your own site and eyeballing it. You catch it by checking what a crawler actually receives: `curl -I` the URL, read the status line, or crawl a batch of nonsense slugs and see what comes back. If the answer is 200 for URLs that shouldn't exist, you have a soft-404 surface, and it is probably bigger than the one page you happened to check.

08

Bottom line

A soft 404 costs you nothing you can see and something real you can't: crawl budget spent on junk, index space occupied by duplicates, and a signal to Google that your domain manufactures thin pages, which is the exact signal a site recovering from a content demotion cannot afford to send. The fix here was one nginx line, already proven correct on three other sections of the same site. Finding it took an actual `curl` against a slug I invented, not a look at the page. If you run a section with any kind of fallback route, one on a React SPA, a CMS, a docs site, go test a slug that does not exist and read the status code back. If it says 200, you have the same bug I had.

---

Written by Andrew Pyle. I build and run this site's own infrastructure, including the prerender and serving layer behind /writing, and I run its SEO recovery as reversible, human-approved agent commands rather than by hand.

Have something you need built or fixed?

I build production Django / Next.js platforms and human-supervised AI-agent systems. Solo, senior, and fast. Tell me what you are building.

Start a project