HEINOUX Journal Portfolio

Journal

I nearly published a true number that proved the opposite

I went to fetch the citation, not to check it. The number was correct and it is still in the paper. The test that produced it was removed over a year ago, and the argument it supported had reversed.

By · · 5 min read

I nearly published a true number that proved the opposite

I had the statistic before I had the post. That should have been the first warning.

The slot for today was about getting a straight answer out of your own numbers, and I knew exactly how I wanted to open it. There is a benchmark called Spider 2.0 that tests whether software can answer a question about a real business database, the kind with thousands of columns and years of accumulated mess in it. When it came out, the results were famously bad. The best model of the day solved something like seventeen percent of the tasks. I had carried that number around for a while because it was useful. It made a point I believe in: the demo works, your actual data does not.

So I went to fetch the citation. Not to check the number, I want to be honest about that. To fetch it. I was confirming a spelling, not testing a claim.

The top of the leaderboard now reads 96.70 percent, on the track where the database comes with prepared metadata and documentation.

The number was never wrong

This is the part I keep turning over. I did not misremember it. I did not read a summary of a summary, which is the mistake I wrote about when I published a figure several sources agreed on and no primary document supported. I did not read the wrong page of the right site, which is the mistake I made with Meta's pricing docs. The number was 17.1 percent. It was published in a real paper in November 2024. It is still in that paper. If you look it up right now you will find it exactly where I remembered it being.

It just does not mean what it meant anymore.

Then it got worse, in the way that makes a thing interesting. I went looking for what the same test scores today, so I could write an honest before and after. There is no after. On 22 May 2025 the people who run the benchmark removed the original setting altogether and replaced it with a different, harder one. The test that produced my number does not exist. Nobody retired it because it was wrong. They retired it because the field had moved past what it was built to measure.

So the number is not merely out of date. It is a score from a race that is no longer run.

That is a different kind of failure and I did not have a name for it. The other two were errors at the moment of reading. This one was correct at the moment of reading and went stale underneath me while I was not looking. There is no source check that catches it, because the source is fine. The source is doing its job. The world moved.

What it would have cost

I would have opened a post with a true sentence that produced a false impression, which is the worst kind of wrong because nobody can point at the error. Every fact in it would have survived a fact check. The whole argument would have been backwards.

And the argument mattered. I was going to tell owners that software cannot yet answer questions about their real data, so do not bother. What the leaderboard actually says is nearly the opposite: on prepared, documented data it now answers almost everything.

I would have talked a reader out of something that works.

The finding was better than the stat

Here is the part that made the day worth it. When I stopped being annoyed and actually read the three tracks, the interesting thing was not the top score. It was the spread.

Ninety six point seven percent on the track where the database has prepared metadata and documentation. Seventy six on the mixed track. Sixty five point six on the one where the system has to go and read a real project and hold the whole context in its head. Same class of technology, roughly thirty one points of difference, and the only thing that changed was how much of a mess the data was in.

I have been saying a version of that to clients for two years without a number behind it. Capture the fact once. Do not keep three copies of one sale. The reason I say it is not tidiness, it is that everything downstream gets easier, and I could never point at evidence. Now I can point at a benchmark where the identical technology loses a third of its accuracy purely because the environment underneath it is untidy.

The stat I went to fetch would have made a small correct point. The stat I found makes a bigger one, and it happens to be the point my whole business rests on.

What I have changed

I have added a rule for myself, and it is narrower than the ones I already had.

Any number in a draft that is more than twelve months old gets re-fetched from the primary source before it ships, even when I am certain of it. Especially when I am certain of it. Certainty is what stopped me checking this morning, and certainty is what the rule has to override.

I nearly did not check today. I want that written down somewhere I will see it again. The reason I checked was administrative, not sceptical. I wanted the spelling of the benchmark right. If it had been spelled the way I assumed, I would have published the opposite of the truth with a completely correct footnote.

The version of this that scares me is not the one I caught. It is the one where the citation was easy to type.

Want this kind of build for your business?

I build AI systems, custom company dashboards and automation for growing businesses.

Get your autopsy Email me