Built means exercised: why a green deploy isn't done
A related-essays card compiled, passed CI, and deployed green. It rendered nothing on all 52 pages. Here is what that taught me about the actual definition of done.
By Andrew Pyle
A feature that compiles, passes CI, and deploys green can still do nothing. I know because I shipped one.
01
A related-essays card that showed nothing
Every essay on this site ends with a "Related" card, links to a few other pieces worth reading next. I built it the normal way. A relationship model in the database connects one essay to others. A serializer walks that relationship and emits the related links as JSON. A React component reads that JSON and renders the card.
Three layers, each one doing its job in isolation, wired end to end. The pull request touched a migration, a serializer field, and a component. It compiled. The test suite passed. Buildkite went green. I merged it and moved on.
The card was empty on all 52 essays.
02
Green all the way down
Nothing in the pipeline objected. The migration ran clean. The serializer had no bug that a unit test would catch, because a unit test would have to know to check for related essays that never existed. The component rendered exactly what it was told to render, which was an empty list. Every gate the deploy process runs said the same thing: fine, ship it.
That is the trap. A green pipeline proves the code compiled and the process completed. It does not prove the feature does its job for a real reader. Those are different claims, and it is easy to build a system that only ever checks the first one.
03
Three ways to build a feature that does nothing
I found what was actually wrong only by looking at the served HTML, and there were three separate things wrong, stacked.
First, nothing had ever populated the relationship the whole feature depended on. The model existed. The column existed. No process, migration, or command had ever written a row into it. The card was empty because there was, quite literally, nothing to show.
Second, even the code that would have emitted a link had the wrong URL baked into it. It built links as `/blog/<slug>` instead of `/writing/<slug>`, an old prefix left over from before the section was renamed. Had the relationship been populated, the card would have rendered links that 404'd.
Third, a separate filter in the API excluded 16 of the 52 essays from the public feed entirely, for reasons that had nothing to do with this feature. Even a correct related-essays query run against the public API would have quietly lost a third of the candidate pool.
Any one of these would have sunk the feature. Together they are the kind of failure a code review does not catch, because each layer looks locally correct. The bug is not in any one file. It is in the gap between the files, and the only place that gap is visible is the actual page.
04
The tell was in the served HTML
I almost missed it. I called populating the relationship a "quick win," wrote the data migration, merged it, and moved to the next task. What stopped me was a habit, not a test: I opened the live page as Googlebot would see it and looked for the specific related slugs I expected to find.
They were not there. Not "there but styled wrong." Not "there but slow to load." Absent. The HTML that a crawler receives, and the HTML I was staring at, had no trace of the feature I had just shipped, reviewed, and deployed.
That is the only way I would have caught it. Reading the code again would not have surfaced it, because the code, read in isolation, looked fine. The test suite would not have surfaced it, because the tests exercised the pieces I remembered to write tests for, and I hadn't thought to assert on the one thing that mattered: that a real essay page, fetched over the wire, contained real related links.
05
Built means exercised
This is item 7 in the platform doctrine I run my agent fleet against: built means exercised. Definition of done is observed behavior in production, not a green deploy.
A green pipeline is a claim about the process. Observed behavior is a claim about the outcome. I had spent years treating the first as a stand-in for the second, because most of the time they move together closely enough that the gap doesn't bite. This time it did, on every single page the feature was supposed to touch.
The fix for how I work was not a better test, though better tests help. It was a rule: nothing counts as done until I have looked at the actual thing a real user, or a real crawler, receives. Not the admin panel. Not the local dev server with seeded data that hides the real gaps. The production response.
06
The fix that worked was the one I watched work
The related-essays feature I described above never got fixed by patching those three bugs. I tried a different approach instead: inline links, placed by hand or by an agent directly into the body of each essay's prose, using a mechanism I could verify by loading the page and clicking the link myself.
I only trusted that approach after doing to it exactly what I should have done to the first one: I loaded a published essay, found the inline link, clicked it, and watched it land on the right page. Not "the code that generates the link looks right." The link, clicked, arriving somewhere real. That is a lower bar architecturally, an inline link is a much simpler thing than a relationship model plus a serializer plus a component, and the simplicity is part of the lesson. The fancier design failed silently. The plain one, checked by hand, worked.
07
Same story, different agent
I run a fleet of agents that build and operate my sites now, and this same failure mode shows up constantly, just with a different narrator. An agent will report "done, deployed, green" with complete confidence, because from where it sits, that is true. It ran the command. The command succeeded. Nothing in its view of the world says otherwise.
That report is not done. It is a claim about the process, made by the part of the system that only sees the process. The way you catch the related-essays bug, whether a person or an agent shipped it, is the same: go look at the real path. Fetch the live page as the crawler fetches it. Query the production database for the row you expect to exist. Click the link. If the agent's job was to make something appear, the definition of done is that it appears, observed, not that the agent said it would.
I have started asking agents to close their own loop this way: after the deploy, go check the actual served output for the specific thing you just built, not a proxy for it. It costs one more step. It would have caught my bug in five minutes instead of after 52 essays shipped with a dead card.
08
Bottom line
I shipped an inert feature. It compiled, it passed CI, it deployed green, and it did nothing on every page it touched, for three unrelated reasons I only found by looking at the live HTML instead of trusting the pipeline. The fix that actually works is not a smarter build system. It is refusing to call anything done until I have watched it work where a real reader, or a real crawler, would see it.
Related
writing
The publishing pipeline that lied to me for four days
A green log, a stale site, and the difference between a job that ran and a job that did the thing.
writing
A worktree per agent, so they never step on each other
The moment you run more than one agent on the same repo, they start clobbering each other's files. The fix isn't cleverer coordination — it's giving each one its own copy of the working tree, so parallelism stops being a fight over shared state.
writing
Agent skills are the reusable unit I was missing
For a year I treated prompts as disposable. Then I started packaging the ones that worked — and my agents stopped relearning the same job every time.
writing
Status is observed, not declared
The most useful thing my command center does is refuse to believe its own status fields — and treat the gap between claimed and true as the actual signal.
writing
Dry-run by default
The one rule that lets me run an autonomous system without lying awake about it: nothing writes to the outside world until I say so — and everything it does, it can undo.