← Writing / June 28, 2026 / 4 min read
Teaching XGBoost to Read Blueprints
Lessons from building a vector-native structural member detector for steel construction drawings, where gradient-boosted trees beat transformers, label quality beat everything, and the headline number turned out to be measuring the wrong thing.
At SteelEye we needed to automate takeoff: given a structural construction drawing, find every placed member (columns, beams, joists, braces, base plates, deck) and assign each its correct AISC section name (W14X22, HSS5X5X1/4, 26KSP…). Estimators do this by hand today, page by page, and it’s slow, error-prone work that million-dollar bids depend on.
Two years of ML hype would tell you to fine-tune a vision-language model and call it a day. Here’s what actually worked, after ~200 experiment scripts and 33k lines of research code.
Work on vectors, not pixels
Construction PDFs aren’t scans; they’re CAD exports. The text tokens and line segments are right there in the file. Instead of rasterizing and running object detection, we extract primitives with PyMuPDF and classify tokens: is this string a member’s name, and if so, what class of member?
This one decision bought us exact text (no OCR noise), exact geometry (segment endpoints, orientations, lengths), and two orders of magnitude less compute than a vision pipeline.
Ground truth came from an unusual place: NC1 (DSTV) files, the CNC instructions the fabrication shop actually cuts from. If the model says a page contains W12X26 beams and the NC1 answer key agrees, that’s validation no human labeling budget could match.
The model zoo, and who survived it
We benchmarked honestly: gradient-boosted trees (LightGBM/XGBoost), a deep residual MLP, a set-attention transformer, a kNN message-passing GNN, tabular foundation models (TabICL, TabPFN), even the Hierarchical Reasoning Model. The numbers below are relative scores from the same self-consistent harness (more on why that harness flattered everyone in a moment), so read them as a ranking, not as capability.
| Model | Relative F1 (same harness) |
|---|---|
| GBM (XGBoost/LightGBM) | best |
| MLP (residual, bf16) | close behind |
| GBM + MLP ensemble | close behind |
| GNN (kNN graph, 3 layers) | worse |
| Transformer (set attention) | diverged |
The gradient-boosted trees won, on 84 engineered features: geometry (distance to segments, orientation context), relational cues (neighbor density, same-row/column), and 32 PCA dimensions of e5-small text embeddings. That last part matters: a 33M-parameter embedding model beat its larger siblings at understanding technical codes, and added +6.6 points on unseen naming schemes.
This mirrors the CAD symbol-spotting literature: state-of-the-art methods there are shallow, kNN-edge-based systems, not deep stacks. When your entities are sparse, structured, and text-anchored, representation beats architecture.
Synthetic data is a representation lever, not a volume lever
The hardest problem was cross-fabricator generalization: every detailing shop has its own drawing style and mark scheme. Real training data all came from one fabricator.
So we built a procedural drawing generator with style personas: synthetic sheets with randomized grids, leader lines, mark conventions, and adversarial decoys (notes and callouts that look like members). Training on real + 500 synthetic sheets lifted novel-fabricator accuracy from 0.57 to 0.78.
Two counterintuitive findings:
- Diversity saturates. 48 personas ≈ 160 personas. Once the model has seen “enough kinds of different,” more variety adds nothing.
- Random mark schemes teach a skill, not facts. Synthetic sheets with nonsense prefixes force the model to learn “unknown prefix → trust the geometry”, exactly the behavior you need on a new fabricator’s drawings.
We also tested the fashionable alternatives: domain-generalization losses (IRM, GroupDRO, CORAL), an RL-based adversarial curriculum, pseudo-labeling. All rejected: none beat plain XGBoost with better data.
Label quality beats model capacity
The single largest jump in the entire program came from a labeling fix, not a modeling idea: marking one positive token per member (its name/mark) instead of every token inside its bounding box. Clean supervision moved metrics more than any architecture change we tried.
If your model has plateaued, audit your labels before reaching for a bigger network. In our callout-detection work, a vision-model audit found 29% of human gold labels were corrupted by an export bug. The “model problem” was a data problem wearing a costume.
Where it landed
For a while I thought the production detector hit 0.987 member-F1 and 0.948 end-to-end, and I reported those numbers. They turned out to be measuring the wrong thing: scored against a benchmark auto-derived from the model’s own inputs, the model was largely grading its own homework. I tell that whole eval-integrity story in The 0.98 F1 That Wasn’t. Regraded against independent human gold, the honest detector reaches ~0.77 member-F1 and ~0.78 end-to-end (found and correctly named) on a frozen human-annotated test set, with leave-one-project-out cross-validation and reproducible-from-scratch docs. That is the number I stand behind, and I think it reads as the stronger result.
The boring stack won: engineered features, gradient-boosted trees, procedural data, obsessive label hygiene. The exciting part isn’t the architecture. It’s that estimators get hours of their week back.
Before a big steel building goes up, someone has to read the drawings and make a list. Every beam, every column, every brace, every base plate: find it, name it, count it. Right now a person called an estimator does this by hand, page by page, and it is slow, tiring work where a single mistake can throw off a bid worth millions of dollars. So the goal of my project was simple to say: teach a computer to do that reading.
Read the file, do not photograph the page
Here is the first clever bit, and it is worth slowing down for.
You might picture a construction drawing as an image, like a photo of a printed page. And if that were true, the computer would have to squint at the picture, guess where the letters are, and try to make out shapes. That guessing is messy and error prone.
But these drawings are not photos. They come out of design software, and inside the file the text is already real text, and the lines are already real lines, with exact positions. Think of the difference between opening a document on your computer, where you can select and copy the words instantly, versus taking a photograph of that same document printed on paper, where the words are now just a blur of pixels you have to decode. One is effortless and exact. The other is a chore full of mistakes.
My tool reads the real text and the real lines straight out of the file. No squinting, no guessing. It gets the exact words and the exact shapes for free, and it does the whole job with a tiny fraction of the computing power a photo based approach would need. That one decision made everything downstream easier.
The old, boring method beat the fancy one
Now for the surprise. If you have heard anything about computers and learning in the last couple of years, you have heard about the big, fashionable systems, the ones that write essays and describe pictures. The obvious move was to reach for one of those.
So I ran a fair contest. I lined up the trendy modern approaches against a much older, plainer family of methods, the kind built out of simple yes or no questions stacked into what people call decision trees. Picture a game of twenty questions: is this line long or short? Is it near the edge of the grid or in the middle? Does it sit next to other similar marks? Ask enough small questions in the right order and you can sort things out remarkably well.
The plain decision tree method won. The flashy modern systems either did worse or fell apart entirely. It turns out that when the thing you are reading is neat, structured, and full of real text already, you do not need a giant brain. You need a tidy set of good questions. Sometimes the boring tool is simply the right tool.
The biggest win came from fixing a dull mistake
This is my favorite part, because it is so unglamorous.
At one point the tool stopped getting better no matter what I tried. The instinct in this field is always to reach for something bigger and smarter. But the single largest improvement in the entire project did not come from a cleverer method at all. It came from fixing a plain, boring labeling mistake in how I was teaching the tool.
When you train a computer like this, you show it examples and tell it the right answers. I had been sloppy. For each steel part I was marking too many bits of text as its name instead of just the one that actually was its name. Once I cleaned that up, so each part had exactly one correct label, the tool leapt forward, further than any fancy method had ever pushed it.
The lesson has stuck with me. When something you built stops improving, do not assume you need a bigger, cleverer engine. Go check whether you have been teaching it wrong. In a related piece of my work, a careful audit found that almost a third of the supposedly correct answers we had been trusting were actually corrupted by a boring software bug. The problem looked like a smart machine problem. It was really a messy homework problem in disguise.
Where it ended up
I also had to solve a harder version of the puzzle: every steel shop draws in its own personal style, and my real examples all came from one shop. To fix that I generated a big batch of fake practice drawings in many invented styles, which taught the tool to trust the shapes when it met an unfamiliar naming style. That lifted its accuracy on a new shop’s drawings a great deal.
In the end the whole thing runs on the plain, unglamorous stuff: good questions, clean labels, careful practice. The exciting part was never the machinery. It is that estimators get hours of their week handed back to them.