On July 1, an eval round on my data agent came back at 32.4%. Two weeks earlier the same suite had scored 66.7%. The factual-accuracy slice — the one that decides whether a data agent is worth deploying at all — had fallen from 65% to 5%.
Two months earlier I would have spent that day tearing the system prompt apart. I know because I have April changelog entries showing me doing exactly that. Instead I followed my process's first rule: verify ground truth before diagnosing. That rule is the output of an improvement loop I've been running since mid-April, and the loop — not the agent's configuration — turned out to be the real thing I was building. This post is the argument for that claim.
First, the story. Twenty minutes of direct queries settled the diagnosis. The test tenant's data had vanished from the current pipeline snapshot. Zero of 2,581 items in the metric store. Zero rows in the reporting views. The eval's answer key described an org that, as far as this snapshot was concerned, no longer existed.
Then I read the traces, and the picture inverted. Across 105 test cases the agent fabricated zero numbers. When the metric-lookup tool returned empty, it pivoted to the SQL tool. When the SQL tool found nothing, it said, plainly, that there was no data for this org. The grader marked those answers wrong because the answer key still had numbers in it. The instrument moved. The agent didn't.
There's a strange bonus buried in this failure. I could not have designed a better honesty stress test than an entire org's data silently disappearing: 105 opportunities to hallucinate a plausible number, zero taken. The round that looked like a catastrophe was the strongest evidence yet that the agent is ready for its beta.
But the part I keep coming back to is the reflex. Checking the environment before blaming the agent is not instinct or talent. It's a written rule that exists because in April I didn't have it and the absence cost me a week. The agent's configuration — prompts, tool schemas, hooks — is not the asset. It's the current output of the asset. The asset is the loop.
What the loop actually is
The agent itself is easy to describe. You ask a plain-English question about a large operational dataset; it either routes to pre-computed metrics or writes SQL against a warehouse. From mid-April to July I evolved its harness — prompts, tools, hooks, context management, evals — with a coding agent doing most of the mechanical work. (The eval side has its own story; see How Do You Know If Your Data Agent Is Any Good? )
The loop is one file and two enforcement docs.
The file is an append-only changelog. Every deliberate change to the harness gets an entry with four parts:
- A triggering signal. A specific trace or a specific metric delta. "It felt wrong" is not a signal and does not get an entry.
- A hypothesized root cause, in one sentence. If I can't state it in one sentence, I don't understand the failure yet.
- Exactly one scoped change. Not "improved the prompt." One rule, one schema field, one hook.
- Verification. The smoking-gun trace re-run, and the full eval suite. Both.
Corrections happen by writing new entries, never by editing old ones. When a change turns out wrong, the record of being wrong stays. This sounds like bureaucratic self-flagellation. It's actually what makes the rest of this post possible: I can compute statistics about my own mistakes because I never erased them.
Two short process docs enforce the rules — the intervention ladder and the changelog discipline I described in The System Prompt Is the Last Resort — and the coding agent loads both automatically, which matters more than the rules themselves. The ladder fires before any change and pushes every fix toward deterministic layers before permitting a new prompt rule; the entry template fires when a change ships. I don't hold the process in my head. The process interrupts me at the two moments it applies.
The ladder earns its keep in the redirections. In May the agent kept presenting numbers from a stale snapshot without saying so. My instinct was a prompt rule: "always state data freshness." The ladder pushed lower. The snapshot date went into the metric tool's response envelope, where the model can't fail to see it, and a hook now flags any final answer that cites a metric without a date. The prompt never changed. The behavior did — and it can't silently regress, because the check is deterministic.
Ten weeks in, the changelog holds 97 entries: 71 changes, 19 observations, 7 verifications. Two changes were reverted. Ten were superseded. Those last numbers are the honest part, and I'll come back to them.
The loop changed its own behavior
The best evidence that the loop works isn't an eval score. It's that the loop's output shifted layers over time, in exactly the direction its own rules predicted.
| Fix layer | April (~43 changes) | May (~14) | June (~14) |
|---|---|---|---|
| Deterministic hooks and guards | 5 | 3 | 8 |
| Tool schemas, error envelopes, retrieval docs | 15 | 7 | 2 |
| System-prompt rules | 6 | 1 | 1 |
(Rows don't sum to the monthly totals; the remainder was eval tooling and infrastructure work.)
April was reactive, per-trace patching. At peak I wrote 25 entries in a single day: read a trace, find the misbehavior, patch the nearest layer, read the next trace. It felt productive, and some of it was. The tool-schema row holds fifteen real fixes — error envelopes that turned silent empty results into explicit "no rows matched" signals the model could act on, retrieval docs that stopped the SQL tool from guessing column names. But the six April prompt rules were each written in an afternoon, against a single trace, and most of what later got reverted or superseded came from exactly that row. A prompt rule that fixes the trace in front of you has roughly even odds against the next fifty traces. A schema constraint mostly holds.
June looks different in every column. One to three entries a day, planned reliability work instead of reaction. The prompt-rule share halved while deterministic guards doubled — the intervention ladder visibly steering. And the single June prompt change shipped the way the April six should have: with a static regression test that fails if the rule is ever deleted, verified against the live deployment the next day.
You can hear the shift in the entry titles. April, roughly: "metric payload lacks units — value rendered as 0.3% in half of runs and 30% in the other half." June, roughly: "eval collapse is a data gap, not an agent regression." The first is me patching an agent that misbehaved. The second isn't about the agent at all.
The diagnosis inverted
Early loop: a trace shows the agent misbehaving, so patch the harness. Late loop: the eval score moved, so prove whether the environment moved first.
Three of the last four apparent regressions turned out to be environment. A staging snapshot arrived stamped 2000-01-01, so every relative-time question — "last quarter," "past 90 days" — correctly returned nothing, and the grader failed the honesty. The reporting views came back empty one week because an upstream job had silently failed to populate them, so the SQL fallback was right-answering against a void. And the vanished tenant data that opened this post.
The principle I'd extract: after about six weeks, the main enemy of an improvement loop is no longer the agent. It's the measurement environment. The mechanism is mundane. The agent's failure modes get engineered away one at a time — that's what the table shows. But the eval stands on pipelines, snapshots, and tenant data that other teams change without telling you. The agent's error rate falls; the instrument's error rate doesn't. The crossover is inevitable. If your habit is still "score dropped, agent got worse," you'll spend days fixing an agent that isn't broken — and worse, your fixes will appear to work for reasons unrelated to the fix, which poisons the changelog with false causality.
The April version of this incident — the first time environment drift masqueraded as a regression — cost a four-entry wild-goose chase: two prompt patches, one tool tweak, and finally the sheepish observation entry noting that the data had been wrong the whole time. Both prompt patches were later superseded. That scar is why "verify ground truth first" is rule one. The loop wrote the rule. The rule saved July 1.
The debt is honest, which is why it needs tracking
I said verification is one of the four required parts of an entry. In practice, 33 of the 71 changes — 46% — shipped with verification deferred: "pending redeploy," "awaiting next eval round." The process permits this, because the realistic alternative isn't more verification. It's people lying in the verification field.
About two-thirds of that debt was later repaid by explicit verification entries — one of the quiet pleasures of the changelog is watching a "pending" from a Tuesday get closed by a verification entry the following week. Twelve changes were never verified at all. The dominant reason is the previous section's villain: the eval environment broke underneath them, rounds stopped being comparable, and the clean before/after they were waiting for never existed. Those twelve are the part of the harness I trust least.
The debt is honest, and that's precisely why it needs tracking rather than shame. A loop that punished deferred verification would simply produce falsified verification.
But there's a sharper meta-lesson. Every remaining weakness in this system is a promise enforced by prose — "verify later," "audit the deferred list weekly," "re-run the round when the environment recovers" — and prose-enforced promises are the bottom rung of my own intervention ladder. The loop's core principle applies to the loop itself, and I haven't finished applying it. A hook that blocks new changes while unverified entries exceed a threshold would turn the promise into a mechanism. That's the next entry.
What carries over
The loop paid for itself in the only currency that matters here: time to a correct diagnosis for a repeated failure class. In April, environment drift disguised as an agent regression cost four entries and most of a week. In June, the same class — the 2000-01-01 snapshot — cost one session. On July 1 it cost one session again, and this time the agent came out exonerated with evidence attached: zero fabrications across 105 cases, against an empty world.
Almost nothing I was proud of in April survived to July unchanged. The prompts were rewritten. The tool schemas grew error envelopes and lost fields. Ten of the seventy-one changes were superseded by better ones. What survived — what compounds — is the changelog, the evidence rules, and the eval. When the next model generation lands, the prompts and tool schemas will be renegotiated in weeks, and that's fine. They were always the output.
The agent is the output. The loop is the product.