← Writing / July 9, 2026 / 5 min read
The 0.98 F1 That Wasn't: Catching Your Own Benchmark Cheating
An eval-integrity story: how a 0.987 member-F1 turned out to be a model grading its own homework, and the general rule for measuring the thing you think you are measuring.
I once shipped a 0.987. Then I proved it was fake. This is the story of how, and why the number I kept instead, roughly 0.77, is the one I am proud of.
At SteelEye the model detects structural members on CAD construction drawings and assigns each its correct AISC section name. It is boosted trees on geometry features plus a small text embedding of each token. How I built that is a separate story, told in Teaching XGBoost to Read Blueprints. This post is not about the build. It is about the evaluation, and about the specific, seductive way a benchmark can lie to you in your own favor.
Two sources of truth, and the trap between them
Ground truth arrived in two forms.
The first was an auto-derived benchmark. I took the NC1 fabrication files, the CNC instructions the shop actually cuts from, and applied their rule logic to the drawing tokens to derive labels programmatically. Cheap, scalable, no human in the loop.
The second was independent human annotation in Label Studio: people looking at drawings and marking members by hand. Slow, expensive, small.
Scored against the auto-derived benchmark, the model looked spectacular: 0.948 end-to-end, 0.987 member-F1. State of the art by any reading. I believed it for longer than I would like to admit.
The catch
Here is what I missed. The auto-derived labels were a near-deterministic function of the same input features the model sees. The rule that generated the “ground truth” and the model being graded were both reading the same drawing tokens, and the rule was simple enough that the model could largely re-learn it.
So the model was not being tested. It was reproducing a labeling rule, and then being graded by that same rule. It was grading its own homework, and of course it got an A.
The tell was there the whole time. A suspiciously high number is a smell, not a trophy. 0.987 on a genuinely hard perception task should have made me suspicious, not proud. Instead it made me stop looking.
Regrading against a truth the model never touched
The fix was to score the exact same model against the independent human annotations, a source of truth with no causal connection to the feature pipeline.
The same model that scored 0.987 against the auto-derived benchmark scored a member-F1 of about 0.25 against human gold.
That is not a small correction. That is nearly the entire result evaporating. The honest, rebuilt model now reaches around 0.77 member-F1 and 0.78 end-to-end on held-out human gold. The written conclusion in the repo is blunt, and I stand by it:
The 0.98 was an illusion of a self-consistent, easier benchmark. Don’t chase 0.98.
What actually went wrong, stated generally
Strip out the steel and the failure is universal. Any ML engineer can walk into it.
A benchmark derived from your own inputs measures label-rule reproduction, not capability. If the process that produces your labels reads the same features your model reads, and that process is learnable, your model will learn it and your benchmark will applaud. You have built a closed loop and called it an evaluation. The score is real; it is just measuring the wrong thing.
The defenses are simple to state and easy to skip under deadline:
- Always hold out a source of truth that is independent of your feature pipeline. If your labels and your model both descend from the same inputs, you have no evaluation, you have a mirror. Human annotation, a different sensor, a downstream physical outcome: something the model cannot have reverse-engineered.
- Treat a suspiciously high number as a smell. The moment a hard task returns an easy score, stop and ask what shortcut you accidentally graded.
- Report the number you can defend to a skeptic. Not the number that looks best in a deck. Imagine someone hostile and competent asking “how do you know that isn’t circular?” and report the number that survives the question. My defensible number is 0.78, and I would rather ship a real 0.78 than a benchmark’s 0.98.
The same mistake wears many costumes
This is one instance of a broader failure mode: measure the thing you think you are measuring, not a proxy that happens to be lying nearby.
I hit the same class of bug from a completely different direction on the speech-recognition side of my work. An early ASR preview looked mysteriously weak, and the cause was that it had been trained on one set of languages and evaluated on a different, non-overlapping set. The evaluation was accidentally zero-shot: the model was being graded on languages it had never seen. Different domain, different symptom, same root cause. The train and test sets did not describe the same thing, so the number was answering a question I had not asked.
In the steel case the eval was too easy because it was circular. In the ASR case the eval was too hard because it was disjoint. Both times the number was confident and both times it was wrong, and both times the fix was the same discipline: go find out, concretely, what your metric is actually a function of.
The rule
A benchmark is a claim about the world, and like any claim it can be self-serving. The auto-derived one flattered me because I built it from the same clay as the model. The honest one, human gold the pipeline never touched, told me the truth, and the truth was a worse number and a better result.
If your evaluation and your model share ancestry, you do not have an evaluation. And if a hard problem hands you an easy score, the burden is on you to prove it is not measuring itself.
I once built something and it scored an almost perfect 98 out of 100. Then I figured out the test was rigged in my own favor, quietly proved it, and watched the number fall to about 77. This is the story of why I am prouder of the 77.
The thing I built reads steel construction drawings, finds every metal part, and gives each one its correct name. How I built it is a separate story. This one is not about the building. It is about the testing, and about a sneaky way a test can flatter you until you believe a comfortable lie.
Two answer keys
To find out how well the tool was doing, I needed a set of correct answers to grade it against. I had two.
The first answer key was cheap and easy to make. There was a set of manufacturing files, the instructions a metal shop uses to actually cut the steel, and I wrote a little rule that turned those files into answers automatically. No people needed, fast, and I could make as many as I wanted.
The second answer key was slow and expensive. It was made by actual humans, sitting down with the drawings and marking each part by hand. Careful, but small.
Graded against the cheap automatic answer key, my tool looked spectacular. Ninety eight out of a hundred. A dazzling score by any standard. I believed it for longer than I would like to admit.
The catch I missed
Here is what I had not noticed, and it is the whole point of the story.
The cheap answer key and the tool were reading the very same drawings, using very nearly the same logic. The rule I used to generate the “correct answers” was simple enough that the tool had basically already learned that exact same rule on its own. So when I graded the tool against those answers, I was not really testing it against reality. I was checking whether it agreed with itself.
Think about a student who is allowed to write the answer key first and then take the test from it. Of course they get an A. The A tells you nothing about whether they understand the subject. It only tells you the two documents came from the same hand.
That is what I had built without realizing it: a loop where the tool graded its own homework against a key it had, in effect, written. And a mirror always tells you that you look great.
The warning sign was there the whole time
The clue was sitting in plain view, and I want to name it because it is the most useful thing here. The task I had set out to do is genuinely hard. Getting 98 out of 100 on a genuinely hard task should have made me suspicious, not delighted. A suspiciously high score is not a trophy. It is a smell. It is the moment to stop and ask what shortcut you accidentally rewarded.
Instead, the wonderful number made me stop looking, which is exactly what a wonderful number does if you let it.
Grading against a real answer key
The fix was straightforward once I saw the problem. I took the exact same tool and graded it again, this time against the slow, expensive answer key that humans had made by hand. That key had no secret connection to my tool. It was made by people looking at the actual drawings, so the tool could not have quietly memorized it.
The score fell hard. The honest, rebuilt version now lands at about 77 out of 100. That is a much less impressive headline, and it is the number I trust completely. I would rather stand behind a real 77 than wave around a fake 98.
The rule, for anyone, not just for computers
Strip away the steel and the drawings and the lesson is simple, and it reaches into everyday life.
If your test and the thing being tested come from the same source, you do not have a test. You have a mirror. A company that grades its own service by asking its own staff whether they did a good job, a study designed by the very people who want a certain result, a quiz written by the same person taking it: all mirrors, all flattering, all useless as real measurement. The fix is always the same. Find a source of truth the thing being tested could not have shaped, and grade against that.
And keep the other half of the lesson close too. When a hard problem hands you an easy, glowing score, do not celebrate yet. Treat that shine as a question, not an answer. Go find out, concretely and honestly, what your test is really measuring. More often than you would guess, the comfortable number is answering a question you never actually asked.
The honest answer key told me a smaller number and a truer one. And a real 77 you can defend to a skeptic beats a beautiful 98 that dissolves the moment anyone leans on it.