← Writing AI · 2026·10·01 · 6 min
FR 00 · Title cardAI

Append-Only or It Didn't Happen

The four changelog entries I wrote while I was wrong are the point. An agent-improvement log has to be non-retconnable.

6 min · Reading mode available:

I recently wrote about the 0.3% bug that wasn't: I spent days fixing an agent under a diagnosis that turned out to be inverted — the "wrong" answer, 0.3%, was correct, and my conviction that it should be 30% came from test fixtures holding a different dataset at a different scale. What I left out of that post is what happened to the four changelog entries I wrote while I was wrong. This post is about those entries, and the argument they make: an agent-improvement changelog has to be non-retconnable. Wrong entries are evidence, not embarrassments. Being able to see what past-you believed, in its original wrong form, is worth more than a clean history.

The agent answers plain-English questions over a large operational dataset, and since April I've been evolving its harness — prompts, tools, evals — with a coding agent doing most of the typing. Every change lands in an append-only changelog — ninety-seven entries in ten weeks, beta launch a few weeks out — governed by one rule: entries are never edited, softened, or deleted after the fact; if an entry turns out to be wrong, you write a new entry saying so.

Four entries deep

The short version, for anyone who skipped the first post. In late April the agent rendered a progress metric as "0.3%" in two of four test runs and "30%" in the other two, on the same stored data. I diagnosed unit ambiguity and spent two days shipping fixes: a self-describing {value, unit, display} envelope around every numeric tool payload, a reorder putting display first, a rendering-guidance section in the agent's domain guide. Four entries deep — observation, envelope, reorder, guide section — each with a confident diagnosis and passing tests, all resting on a premise nobody had checked. When the agent then answered "0.3%" in every run, I logged it as total failure.

Then I read the actual row in the live store. It held 0.3 — mathematically correct across 7,444 records with near-zero engagement. My "30" came from fixtures where the same field held values in the fifties and sixties: different dataset, different scale. The "30%" renders had been the agent hallucinating a factor of 100 onto a bare float — and the fixes had stopped it. Two hallucinated runs in four before; zero in four after. Real hardening, real effect, shipped under a phantom diagnosis and scored against an inverted baseline that made perfect output read as total failure.

What append-only did with the wreckage

Here is where the log's one rule earned its keep. The four wrong entries were not edited. They were not softened with a retroactive "note: this turned out to be wrong" spliced inline. They were not deleted. They read today exactly as they read in April, complete with confident diagnosis and passing-test evidence.

Instead, a new observation entry went in. It says, roughly: the originating premise was false. The live value is 0.3, and "0.3%" was the right answer all along; nothing in the pipeline multiplies or divides by 100, so the unit-ambiguity diagnosis was a phantom. But the fixes stay, reclassified from bugfix to defensive hardening — because they demonstrably stopped a real failure, the model inventing a scale for a bare float, and because the next dataset we onboard may genuinely store the same field at a different scale; now the payload says so explicitly. Wrong reason, right change, measured effect. The entry records all three, plus the part I'd most like to sand off: that I judged the result against an inverted baseline and spent a day reading 100% accuracy as 100% failure.

The only edit permitted to the old entries was a one-line status flip: "superseded by [the correction]." A pointer, not a rewrite. The original reasoning stays intact as the record of what we believed and why we believed it.

Consider the alternative. If I had quietly rewritten those four entries — merged them into one tidy "added unit envelope (hardening)" note — the log would look cleaner, and it would lie. It would erase the most instructive artifact the log has ever produced: a timestamped specimen of a plausible diagnosis, two effective fixes, a false premise underneath, and an evaluator who couldn't see the fixes working because he was sure of the wrong answer. I had already retconned the story once, in my head, when I scored perfect output as failure. The log is the thing that un-retconned it.

The rules that make it work

Append-only alone isn't enough; an append-only pile of vague notes is just a pile. Four other rules do the load-bearing work.

One change per entry. When an entry bundles a prompt tweak, a tool fix, and a threshold change, and the eval moves, you cannot attribute the delta. Unattributable deltas are the root cause of silent regressions — you keep the whole bundle because "it helped," including the part that hurt. This rule is why the phantom chain was four entries instead of two: the envelope, the reorder, and the guide section were separate changes with separate entries, so each could later be judged on its own. It is also how the envelope survived the correction — it had its own entry and its own measurable effect, so it could be re-scored instead of bulk-reverted along with the diagnosis.

Every entry names its triggering signal. A specific trace, a specific metric delta, a specific row. Never "responses felt worse this week." The phantom entries followed this rule — they named concrete renders in concrete runs — which is exactly why the correction entry could later pin down where the observation went wrong. A vague trigger can't be falsified, and an entry that can't be falsified is dead weight.

Status is explicit: shipped, reverted, or superseded-by. The most common way an agent harness rots, in my experience, is that a reverted fix gets quietly re-introduced three weeks later by someone — often me — who forgot why it was removed. The log can only prevent that if reverts are first-class and visible, not scrubbed.

Corrections are new entries plus one pointer flip. Never edits. This is the rule the other three exist to serve.

The burn became a rule, and the rule fired

The published post ends with a lesson: verify the "wrong" answer against live data before diagnosing the agent. The correction entry did one more thing with it — it turned the sentence into procedure. Step 0 on the pre-change checklist: if your triggering signal names a specific value, row, or config, fetch it from the live source before you reason about it. A patch chain built on an unverified data claim is worse than no patch, because each new link manufactures fake confirmation for the previous one. Learning that lesson was the first post. Enforcing it is this one.

Step 0 has fired twice since.

In June, a query category started failing, and three separate code-reading passes — mine, the coding agent's, a colleague's — converged on the same diagnosis: an ID format mismatch between two layers. Unanimous, coherent, wrong. Step 0 demanded one live count query before any fix shipped. The count came back zero: the upstream views were empty. No IDs were mismatching because no IDs were arriving. A single query killed a diagnosis that three readers had confirmed from the code alone.

A few weeks later, days before the beta cut, an eval pass rate collapsed from 67% to 32% between runs. In April I would have opened the prompt file. Instead, one live query against the test tenant showed its data had vanished from the pipeline snapshot — the agent was correctly reporting that there was nothing to report. Reading the failing transcripts confirmed it: across all 105 cases, zero fabricated numbers. The agent was exonerated before a single "fix" went in.

Same failure class every time: a confident diagnosis resting on an unverified claim about live data. In April it cost four entries and several days. In June it cost one query and one session. That delta is the changelog working — not as a record of what changed, but as the mechanism that converts one expensive mistake into a rule that pre-empts the next three.

Isn't the wreckage just noise?

The obvious objection: ninety-seven entries, four of them wrong and prominently so — doesn't keeping the wrong ones around clutter the log? Shouldn't it be a clean reference of what's true now?

No. A clean reference of current truth is what the code and the domain guide are for. The changelog's job is different: it is the record of how we got here, and every question I actually ask of it is a question about past states, including wrong ones. Why does this envelope exist? Has this fix been tried and reverted before? What did we believe when we shipped that? Sand off the wrong states and the log answers none of them.

The phantom chain is the sequence I reread most. It is the log's best teaching artifact, and its best argument for itself.

So: when a past decision looks wrong in hindsight, say so — in a new entry. The wrongness is not a blemish on the record. It is the record. A changelog you can retcon isn't a log; it's a press release.

← Previous · 2026·09·24 The Outer Loop Is the Product All writing → This is the latest. Everything else is in the index.
Reply hi [at] huanyu.io Home