All writing
Notes on building

The Placebo Line

Three months building a research library, one week measuring it. The gate I trusted scored zero, gibberish tied my proudest number, and the only result worth keeping was a null.

August 8, 2026 · 5 min

For the last three months, a growing share of my machine's day has gone into one asset: a research library. Agents read source material the way a new hire reads the operations binder, then file what they learn as searchable passages. There are just over thirty thousand of those passages now, spread across twenty-three libraries, one per subject. When an agent takes on real work, it searches the shelves before it reasons. The model rents me intelligence by the minute. The library is supposed to be the part I own.

This week I stopped building it and measured it. Not spot checks. Written questions with known answers, scored against the shelves, library by library, all of it at once, for the first time. Four findings came back. Every one was embarrassing, every one was cheap to have prevented, and I suspect most people assembling a library like this right now are carrying all four without knowing it. Here is the week, in the order it hurt.

The gate that scored zero

I had a quality gate. Its job was to curate the question bank, keeping the questions worth asking and discarding the rest, and for months I treated what it kept as my best material. When the measurement finally ran, the questions the gate had kept scored zero. Not low. Zero. The gate had been selecting, with perfect consistency, for exactly the questions that cannot find their answers in the library.

How does a backwards instrument survive for months? Because nothing in its life ever forced it to be right. It ran on schedule. It kept some things and rejected others, which felt like judgment. Its output was tidy and looked like rigor. But its output fed no scoreboard, so the one property that mattered, whether its keep pile was any good, was the one property nothing ever checked. An instrument that cannot fail is a label, not a measure. The gate could not fail. So it never did, and it was wrong the entire time. It is gone now, replaced by one whose keep pile gets scored. Replaced, in other words, by one that can lose.

What nonsense scores

My proudest number went next. One library carried a headline I had quoted more than once: ten of ten, full coverage of the topics that matter. This week a reviewer whose whole brief was to attack that number ran the control I should have run on day one. Take a list of generic words, the kind that appear in every document ever written about the subject, words carrying no knowledge at all, and score the same claim with them. The nonsense went ten for ten. My headline was reproduced, in full, by a placebo.

Ten coverage claims against the score that nonsense earns Ten horizontal bars, one per coverage claim. A dashed gold vertical line marks what a placebo of generic words scores on the same test. Eight faint bars stop at or below the line and only two solid bars clear it. Only what clears the line is a finding. COVERAGE VS THE PLACEBO LINE what nonsense scores Only what clears the line is a finding. /ar/

The placebo is permanent now. Every coverage claim in the library prints beside the score that generic filler earns on the same test, and nothing counts as a finding unless it clearly beats the filler. The first time I drew that line on the chart, most of my bars did not clear it. The board looked worse than it had ever looked, and it was more true than it had ever been.

A copy of a reading is not a second reading

The library does not hold sources alone. Agents also write summaries, the condensed briefings other agents actually read first, and every claim in a summary carries a citation back to the passage it came from. Those citations were why I trusted the summaries. This week I audited the load-bearing claims by doing the one thing a citation invites and almost nobody does. I opened them. Two hundred and five claims, each checked against the exact passage it cites. One hundred and seven failed. Just over half. Claims that said more than their source says. Claims that dropped a hedge the source was careful to keep. Claims that were true about an old version and stated as timeless.

The citations had reassured me for a bad reason. Each citation was written in the same pass, by the same reader, as the claim it supports. Checking a claim against its own citation feels like a second opinion, but it is the first opinion, photocopied. A copy of a reading is not a second reading. Verification starts when different eyes open the source with no stake in the claim surviving. All one hundred and seven failures now carry corrections in place, so the next reader inherits the audit instead of the error.

The zero I kept

Then the test I expected to enjoy. With the summaries repaired, I re-ran the full measurement and waited for the retrieval scores to rise. They moved by zero point zero. Not one library improved. My first instinct was that the re-run was broken, so I proved, the slow way, that the test was really reading the repaired library and not some stale copy. It was. The zero was real.

That zero is the most valuable thing the week produced, because of what it separates. Retrieval finds documents. The corrections did not change which documents get found. They change what a reader believes after finding one. Those are two different halves of quality. I had been measuring the first half, repairing the second, and waiting for the first needle to thank me for it. It was never going to. The two halves are measured on their own instruments now.

The null that survives verification is worth more than the win that does not. I have shipped wins that evaporated the first time anyone touched them. This zero was checked twice, it will still be true next month, and it says precisely where the next hour of work should go and where it would be wasted. That is more than most of my green numbers have ever done for me.

The inventory

On paper the library had a terrible week. A curated question bank went to the shredder. A ten of ten became a claim that gibberish can tie. Half the load-bearing claims in the summaries now wear corrections. The repair I was proudest of moved the headline metric by exactly nothing. And every line of that is an upgrade, because every number that survived the week is now a number I can hand to someone else without flinching.

I have watched the same disease in the restaurant for years, wearing an apron instead of a dashboard. The checklist that comes back perfect every single shift. The week with no complaints, taken from a floor where nobody asked a single table how anything was. The supplier scorecard filled out by the supplier. None of those instruments can fail, so none of them are instruments. They are wallpaper with numbers printed on it.

Three months of building, and the honest inventory took one week. The week was only possible because, for the first time, the library was allowed to lose. That is the test I would give any system that reports its own quality, in a kitchen or a data center. Do not ask how good its numbers are. Ask when they last had a chance to be bad.

/ar/