<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Irving Ernesto · Writing</title><description>AI engineer &amp; physicist. CTO of RoomIQ, main developer of SteelEye. I build ML systems, data pipelines and the infrastructure that keeps them alive in production.</description><link>https://irvingernesto.com/</link><language>en-us</language><item><title>The 0.98 F1 That Wasn&apos;t: Catching Your Own Benchmark Cheating</title><link>https://irvingernesto.com/blog/the-f1-that-wasnt/</link><guid isPermaLink="true">https://irvingernesto.com/blog/the-f1-that-wasnt/</guid><description>An eval-integrity story: how a 0.987 member-F1 turned out to be a model grading its own homework, and the general rule for measuring the thing you think you are measuring.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;I once shipped a 0.987. Then I proved it was fake. This is the story of how, and why the number I kept instead, roughly 0.77, is the one I am proud of.&lt;/p&gt;
&lt;p&gt;At &lt;a href=&quot;https://steeleye.ai&quot;&gt;SteelEye&lt;/a&gt; the model detects structural members on CAD construction drawings and assigns each its correct AISC section name. It is boosted trees on geometry features plus a small text embedding of each token. How I built that is a separate story, told in &lt;a href=&quot;/blog/teaching-xgboost-to-read-blueprints&quot;&gt;Teaching XGBoost to Read Blueprints&lt;/a&gt;. This post is not about the build. It is about the evaluation, and about the specific, seductive way a benchmark can lie to you in your own favor.&lt;/p&gt;
&lt;h2&gt;Two sources of truth, and the trap between them&lt;/h2&gt;
&lt;p&gt;Ground truth arrived in two forms.&lt;/p&gt;
&lt;p&gt;The first was an &lt;strong&gt;auto-derived benchmark&lt;/strong&gt;. I took the NC1 fabrication files, the CNC instructions the shop actually cuts from, and applied their rule logic to the drawing tokens to derive labels programmatically. Cheap, scalable, no human in the loop.&lt;/p&gt;
&lt;p&gt;The second was &lt;strong&gt;independent human annotation&lt;/strong&gt; in Label Studio: people looking at drawings and marking members by hand. Slow, expensive, small.&lt;/p&gt;
&lt;p&gt;Scored against the auto-derived benchmark, the model looked spectacular: 0.948 end-to-end, 0.987 member-F1. State of the art by any reading. I believed it for longer than I would like to admit.&lt;/p&gt;
&lt;h2&gt;The catch&lt;/h2&gt;
&lt;p&gt;Here is what I missed. The auto-derived labels were a near-deterministic function of the same input features the model sees. The rule that generated the “ground truth” and the model being graded were both reading the same drawing tokens, and the rule was simple enough that the model could largely re-learn it.&lt;/p&gt;
&lt;p&gt;So the model was not being tested. It was reproducing a labeling rule, and then being graded by that same rule. It was grading its own homework, and of course it got an A.&lt;/p&gt;
&lt;p&gt;The tell was there the whole time. &lt;strong&gt;A suspiciously high number is a smell, not a trophy.&lt;/strong&gt; 0.987 on a genuinely hard perception task should have made me suspicious, not proud. Instead it made me stop looking.&lt;/p&gt;
&lt;h2&gt;Regrading against a truth the model never touched&lt;/h2&gt;
&lt;p&gt;The fix was to score the exact same model against the independent human annotations, a source of truth with no causal connection to the feature pipeline.&lt;/p&gt;
&lt;p&gt;The same model that scored 0.987 against the auto-derived benchmark scored a &lt;strong&gt;member-F1 of about 0.25&lt;/strong&gt; against human gold.&lt;/p&gt;
&lt;p&gt;That is not a small correction. That is nearly the entire result evaporating. The honest, rebuilt model now reaches around 0.77 member-F1 and 0.78 end-to-end on held-out human gold. The written conclusion in the repo is blunt, and I stand by it:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The 0.98 was an illusion of a self-consistent, easier benchmark. Don’t chase 0.98.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;What actually went wrong, stated generally&lt;/h2&gt;
&lt;p&gt;Strip out the steel and the failure is universal. Any ML engineer can walk into it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A benchmark derived from your own inputs measures label-rule reproduction, not capability.&lt;/strong&gt; If the process that produces your labels reads the same features your model reads, and that process is learnable, your model will learn it and your benchmark will applaud. You have built a closed loop and called it an evaluation. The score is real; it is just measuring the wrong thing.&lt;/p&gt;
&lt;p&gt;The defenses are simple to state and easy to skip under deadline:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Always hold out a source of truth that is independent of your feature pipeline.&lt;/strong&gt; If your labels and your model both descend from the same inputs, you have no evaluation, you have a mirror. Human annotation, a different sensor, a downstream physical outcome: something the model cannot have reverse-engineered.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Treat a suspiciously high number as a smell.&lt;/strong&gt; The moment a hard task returns an easy score, stop and ask what shortcut you accidentally graded.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Report the number you can defend to a skeptic.&lt;/strong&gt; Not the number that looks best in a deck. Imagine someone hostile and competent asking “how do you know that isn’t circular?” and report the number that survives the question. My defensible number is 0.78, and I would rather ship a real 0.78 than a benchmark’s 0.98.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;The same mistake wears many costumes&lt;/h2&gt;
&lt;p&gt;This is one instance of a broader failure mode: &lt;strong&gt;measure the thing you think you are measuring, not a proxy that happens to be lying nearby.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;I hit the same class of bug from a completely different direction on the speech-recognition side of my work. An early ASR preview looked mysteriously weak, and the cause was that it had been trained on one set of languages and evaluated on a different, non-overlapping set. The evaluation was accidentally zero-shot: the model was being graded on languages it had never seen. Different domain, different symptom, same root cause. The train and test sets did not describe the same thing, so the number was answering a question I had not asked.&lt;/p&gt;
&lt;p&gt;In the steel case the eval was too easy because it was circular. In the ASR case the eval was too hard because it was disjoint. Both times the number was confident and both times it was wrong, and both times the fix was the same discipline: go find out, concretely, what your metric is actually a function of.&lt;/p&gt;
&lt;h2&gt;The rule&lt;/h2&gt;
&lt;p&gt;A benchmark is a claim about the world, and like any claim it can be self-serving. The auto-derived one flattered me because I built it from the same clay as the model. The honest one, human gold the pipeline never touched, told me the truth, and the truth was a worse number and a better result.&lt;/p&gt;
&lt;p&gt;If your evaluation and your model share ancestry, you do not have an evaluation. And if a hard problem hands you an easy score, the burden is on you to prove it is not measuring itself.&lt;/p&gt;</content:encoded><category>ML</category><category>Evaluation</category><category>SteelEye</category><category>Benchmarking</category><category>Research</category></item><item><title>GSPO: Reinforcement Learning for Low-Resource Speech Recognition</title><link>https://irvingernesto.com/blog/gspo-reinforcement-learning-for-asr/</link><guid isPermaLink="true">https://irvingernesto.com/blog/gspo-reinforcement-learning-for-asr/</guid><description>Teaching Whisper to hear Mexican Indigenous languages with RL. Hand-rolling GSPO around an audio policy, rewarding on negative Character Error Rate, and making real progress on languages every off-the-shelf system fails.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Every state-of-the-art speech recognizer can hear Spanish. None of them can hear the languages Spanish landed on top of. Point Whisper, NVIDIA Parakeet or Canary, Meta MMS, or IBM Granite at Mexican Spanish and you get roughly 14% word error rate: usable. Point the exact same models at Nahuatl, or any of the 22 other Indigenous languages of Mexico, and word error rate jumps to 99% or worse. The model is not transcribing. It is guessing in the wrong language.&lt;/p&gt;
&lt;p&gt;That gap, 14% versus 99%, is the whole project. It is not a tuning problem. It is a “the training data never existed” problem, spanning 23 languages across six or more language families. This post is about closing part of that gap with reinforcement learning, and about the several ways I nearly fooled myself while doing it.&lt;/p&gt;
&lt;h2&gt;First you need a benchmark you cannot cheat&lt;/h2&gt;
&lt;p&gt;Before touching a model I built the evaluation, because in low-resource work the benchmark is where projects quietly die. If you cannot tell progress from noise, every training run looks like progress.&lt;/p&gt;
&lt;p&gt;The result is MEXA, a contamination-resistant ASR benchmark. Two pieces make it hard to game:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;private held-out test set&lt;/strong&gt; that never ships.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;public audio-plus-text fingerprint registry&lt;/strong&gt;, so anyone can decontaminate their own training data against the test without ever seeing the test itself. You check whether your clips collide with the fingerprints; you learn nothing about the answers.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The scoring is linguistically aware. A naive normalizer would strip the tone marks and glottal stops that carry meaning in these languages, quietly rewarding a model for mangling them. MEXA’s normalization preserves those marks through both WER and CER, so the metric measures the language people actually speak, not a flattened ASCII shadow of it.&lt;/p&gt;
&lt;h2&gt;Supervised first, then reinforcement&lt;/h2&gt;
&lt;p&gt;The base model is &lt;code&gt;openai/whisper-large-v3-turbo&lt;/code&gt;. The pipeline is two stages.&lt;/p&gt;
&lt;p&gt;Stage one is a supervised fine-tune: LoRA through Unsloth, with the hyperparameters found by Optuna rather than by me guessing. The winning configuration used rsLoRA plus DoRA at rank 128. This gets the model into the neighborhood of the languages. It stops guessing in Spanish and starts producing plausible orthography.&lt;/p&gt;
&lt;p&gt;Stage two is where it gets interesting: reinforcement learning with GSPO, Group Sequence Policy Optimization (&lt;a href=&quot;https://arxiv.org/abs/2507.18071&quot;&gt;arXiv 2507.18071&lt;/a&gt;).&lt;/p&gt;
&lt;h2&gt;Why I had to hand-roll the trainer&lt;/h2&gt;
&lt;p&gt;The obvious move is to reach for TRL’s &lt;code&gt;GRPOTrainer&lt;/code&gt; and be done in an afternoon. It does not work, and the reason is structural: every mainstream RL-for-LLM library assumes text in, text out. The policy is a language model, the prompt is tokens, the completion is tokens.&lt;/p&gt;
&lt;p&gt;Whisper’s input is not tokens. It is audio. The policy conditions on a log-mel spectrogram, samples a transcript, and gets scored against a reference. So the entire trainer had to be rebuilt around an &lt;strong&gt;audio policy&lt;/strong&gt;: batching audio, generating groups of candidate transcripts per clip, computing rewards, and pushing the GSPO update, with the encoder in the loop the whole way.&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;for batch of audio clips:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;for each clip:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;        &lt;/span&gt;&lt;/span&gt;&lt;span&gt;sample G candidate transcripts from the policy   # audio -&amp;gt; text&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;        &lt;/span&gt;&lt;/span&gt;&lt;span&gt;reward[g] = -CER(candidate[g], reference)        # verifiable, no reward model&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;advantages = group_normalize(rewards)                # per-clip baseline&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;ratio = exp((logp_theta - logp_old) / len(y))        # sequence-level, length-normalized&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;loss = -(ratio * advantages) + beta * KL(policy || sft_anchor)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;step(loss)                                            # lr = 1e-6&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;div&gt;&lt;div&gt;&lt;/div&gt;&lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;h2&gt;The reward is just negative CER, and that is the point&lt;/h2&gt;
&lt;p&gt;The reward function is negative Character Error Rate against the reference transcript. That is it. No learned reward model, no human preference data. This is RLVR, reinforcement learning from a &lt;strong&gt;verifiable&lt;/strong&gt; reward: the reference is ground truth, so the reward is exact and free of the reward-hacking pathologies that plague learned critics.&lt;/p&gt;
&lt;p&gt;It also has an elegant side effect. Whisper’s worst failure mode on hard audio is hallucination: repetition loops where it emits the same phrase over and over. Under a CER reward those loops punish themselves, because every hallucinated character counts as an insertion error and inflates CER. The model is not told “stop looping.” It discovers that looping is expensive.&lt;/p&gt;
&lt;p&gt;To keep the policy from drifting off a cliff, there is a KL penalty (the k3 estimator) back to a frozen SFT anchor, and the learning rate is deliberately tiny at 1e-6. RL here is a scalpel, not a hammer.&lt;/p&gt;
&lt;h2&gt;GSPO versus GRPO, briefly&lt;/h2&gt;
&lt;p&gt;The difference from GRPO is the importance ratio. GRPO computes it per token. GSPO computes it once per &lt;strong&gt;sequence&lt;/strong&gt;, length-normalized:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;ratio = exp( (log p_theta(y|x) - log p_old(y|x)) / |y| )&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;div&gt;&lt;div&gt;&lt;/div&gt;&lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;For long sequences the token-level product in GRPO accumulates variance and can blow up. The sequence-level, length-normalized ratio is far more stable, which matters when your outputs are full transcripts rather than short answers.&lt;/p&gt;
&lt;h2&gt;Results&lt;/h2&gt;
&lt;p&gt;On the 240-clip MEXA subset, evaluated fairly through faster-whisper:&lt;/p&gt;




















&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Stage&lt;/th&gt;&lt;th&gt;WER&lt;/th&gt;&lt;th&gt;CER&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;SFT preview&lt;/td&gt;&lt;td&gt;68.5&lt;/td&gt;&lt;td&gt;30.6&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;SFT + GSPO&lt;/td&gt;&lt;td&gt;66.0&lt;/td&gt;&lt;td&gt;28.3&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;A net improvement of 2.45 WER from RL alone. The comparison that actually matters is not my own benchmark, it is the gap to what already exists: off-the-shelf systems (Whisper, Parakeet, Canary, MMS, Granite) sit near or above 99% WER on these languages, effectively unusable, and where the next-best measured system lands at 74.3 CER, this one reaches 30.6. The point is not a scoreboard rank on a benchmark I built, it is that the languages were near-untranscribable and now they are not.&lt;/p&gt;
&lt;p&gt;The clearest evidence that the CER reward did what it was designed to do: the worst looping language dropped from 100.2 WER to 75.9. That is the hallucination penalty working exactly as predicted, on the language that needed it most.&lt;/p&gt;
&lt;h2&gt;The negative result I am keeping in the paper&lt;/h2&gt;
&lt;p&gt;Here is the part I want to be honest about, because it taught me more than the win did.&lt;/p&gt;
&lt;p&gt;I expected fancier reward shaping to help. Composite rewards, and an MGPO variant that biases toward informative examples. Neither beat plain negative-CER GSPO. Not “beat it by a little,” did not beat it at all.&lt;/p&gt;
&lt;p&gt;The reason is that low-resource ASR difficulty is &lt;strong&gt;bimodal&lt;/strong&gt;, not smooth. The easy Spanish clips have a success probability near 1. The near-impossible Indigenous clips have a success probability near 0. In between there is only a thin frontier of clips where the model sometimes succeeds and sometimes fails, and the frontier is where learning signal lives. When I set a 0.5 difficulty threshold to find those frontier clips, roughly 23 of 50 batches had &lt;strong&gt;zero&lt;/strong&gt; usable frontier clips. The fancier method was starving. It was engineered to exploit a gradient of difficulty that mostly is not there. Plain GSPO, which does not depend on that structure, kept learning.&lt;/p&gt;
&lt;p&gt;The lesson generalizes: sophistication that assumes a smooth difficulty landscape is worse than useless on a bimodal one.&lt;/p&gt;
&lt;h2&gt;Two engineering war stories&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The accidentally zero-shot preview.&lt;/strong&gt; An earlier preview looked oddly weak, and it took a while to see why: it had been trained on one set of languages and evaluated on a different, non-overlapping set. The model was being graded, zero-shot, on languages it had literally never trained on. Lesson, now written down where I will see it: always confirm that your train and test label or language sets actually intersect before you believe any number.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The 30-second window.&lt;/strong&gt; Evaluation clips run around 60 seconds, but Whisper’s attention window is 30. Naively you lose everything past the boundary. The fix is to force-align the transcript to the audio with MMS_FA, then cut on inter-word silences into chunks of at most 28 seconds, preserving the original orthography across the seams so the tone marks and glottal stops survive the split.&lt;/p&gt;
&lt;h2&gt;Where this sits&lt;/h2&gt;
&lt;p&gt;This is part of the broader low-resource line I publish openly on HuggingFace at &lt;a href=&quot;https://huggingface.co/Thermostatic&quot;&gt;Thermostatic&lt;/a&gt;, alongside the translation work. The models get replaced; the benchmark and the datasets are the durable contribution. MEXA exists so that the next person trying to make a speech model hear Nahuatl starts with a number they can trust, and a way to prove they did not cheat to get it.
&lt;/p&gt;</content:encoded><category>ASR</category><category>Reinforcement Learning</category><category>Whisper</category><category>Low-resource</category><category>Nahuatl</category></item><item><title>Reading Refusal Before the Model Speaks</title><link>https://irvingernesto.com/blog/reading-refusal-before-the-model-speaks/</link><guid isPermaLink="true">https://irvingernesto.com/blog/reading-refusal-before-the-model-speaks/</guid><description>An interpretability study with the Jacobian lens: the model commits to refusing roughly ten layers before it writes a word, that decision is legible in the verbalizable workspace, and a surgical pullback edit removes only about a third of it. Where the rest of the &apos;no&apos; actually lives.</description><pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The most surprising thing I found is not that you can edit a language model’s refusal. It is that when you locate refusal precisely, read it cleanly, and surgically remove exactly the thing you were reading, the model keeps refusing anyway. About two thirds of the “no” is somewhere you were not looking.&lt;/p&gt;
&lt;p&gt;This is a small study built on top of Anthropic’s &lt;a href=&quot;https://github.com/anthropics/jacobian-lens&quot;&gt;Jacobian lens&lt;/a&gt;, the reference implementation for &lt;a href=&quot;https://transformer-circuits.pub/2026/workspace/index.html&quot;&gt;&lt;em&gt;Verbalizable Representations Form a Global Workspace in Language Models&lt;/em&gt;&lt;/a&gt;. It is squarely interpretability and safety work: the question is not how to stop a model refusing, it is where and when inside the network that refusal is decided, whether that decision is legible, and how much of it you can actually reach from the part of the model you can read. The headline result is that the readable part is not the load-bearing part, and that is the interesting, safety-relevant finding.&lt;/p&gt;
&lt;h2&gt;Refusal is a computation that finishes before the first token&lt;/h2&gt;
&lt;p&gt;The model here is &lt;code&gt;Qwen3.5-4B&lt;/code&gt;: 32 layers, residual width 2560, with the pre-fitted Hub lens &lt;code&gt;neuronpedia/jacobian-lens @ qwen-n1000&lt;/code&gt;. The Jacobian lens gives an average forward map &lt;span&gt;&lt;span&gt;JlJ_l&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;J&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; from a layer-&lt;span&gt;&lt;span&gt;ll&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; residual to the final logits. That linearization is what lets me ask a causal question about the future: not “what is written in this residual now” but “which directions in this residual push the model toward saying a refusal word later.”&lt;/p&gt;
&lt;p&gt;So I read the J-space at the assistant-generation position, the moment right before the model emits its first token, on harmful versus benign chat prompts. On harmful prompts the workspace lights up &lt;code&gt;Cannot&lt;/code&gt;, &lt;code&gt;cannot&lt;/code&gt;, the Chinese &lt;code&gt;无法&lt;/code&gt;, and &lt;code&gt;illegal&lt;/code&gt; at layers 16 to 24. Benign prompts do not. The relative refusal-mass is roughly +7 for harmful prompts against roughly 0 for benign ones.&lt;/p&gt;
&lt;p&gt;Read that again in terms of time. Nothing has been generated yet. The model has not written “I”. And already, about ten layers deep from where the refusal tokens finally surface, the decision to refuse is present and legible. Refusal is not something the model talks itself into as it writes. It is a computation that has essentially finished before the first token, and the lens lets you watch it finish.&lt;/p&gt;
&lt;h2&gt;The pullback: refusal as it lives in the workspace&lt;/h2&gt;
&lt;p&gt;Here is the mechanical idea. The standard way to remove refusal is &lt;em&gt;abliteration&lt;/em&gt; (Arditi et al. 2024): take the mean difference between harmful and harmless activations, call it the “refusal direction”, and project it out of the residual stream everywhere. It works, but that direction is derived from what correlates with harmful &lt;em&gt;input&lt;/em&gt;, and its damage is only ever checked at the &lt;em&gt;output&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;The Jacobian lens offers a sharper handle. Because &lt;span&gt;&lt;span&gt;JlJ_l&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;J&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; maps a residual to logits, the residual directions that &lt;em&gt;cause&lt;/em&gt; a future refusal token are the pullback of that token’s unembedding through &lt;span&gt;&lt;span&gt;JlJ_l&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;J&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;. Concretely I build&lt;/p&gt;
&lt;p&gt;&lt;span&gt;&lt;span&gt;dl=Jl⊤ (g⊙w),w=mean⁡(W[refusal])−mean⁡(W),g=final-norm gaind_l = J_l^{\top}\,(g \odot w), \qquad w = \operatorname{mean}(W[\text{refusal}]) - \operatorname{mean}(W), \qquad g = \text{final-norm gain}&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;d&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;J&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;⊤&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;g&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;⊙&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;w&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;w&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;mean&lt;/span&gt;&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;W&lt;/span&gt;&lt;span&gt;[&lt;/span&gt;&lt;span&gt;&lt;span&gt;refusal&lt;/span&gt;&lt;/span&gt;&lt;span&gt;])&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;−&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;mean&lt;/span&gt;&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;W&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;g&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;final-norm gain&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;where &lt;span&gt;&lt;span&gt;WW&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;W&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; is the unembedding matrix and &lt;span&gt;&lt;span&gt;ww&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;w&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; is the refusal-token covector, mean-centered. &lt;span&gt;&lt;span&gt;dld_l&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;d&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; is refusal &lt;em&gt;as it lives in the verbalizable workspace&lt;/em&gt;: the per-layer residual direction whose only job is to steer the model toward narrating “I cannot”. No harmful/harmless contrast set is needed to find it, just the pullback of the refusal tokens themselves.&lt;/p&gt;
&lt;p&gt;To edit, I ablate with a reversible forward hook that projects the residual orthogonal to that direction, &lt;span&gt;&lt;span&gt;h′=h−α Q⊤(Qh)h&apos; = h - \alpha\, Q^{\top}(Q h)&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;h&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;′&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;h&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;−&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;α&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;Q&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;⊤&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;Q&lt;/span&gt;&lt;span&gt;h&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;, at every fitted layer from 8 onward, for the duration of one generate pass. Nothing is written to weights. And critically I measure collateral &lt;em&gt;inside the interpretable workspace&lt;/em&gt;: the off-refusal-axis KL of the J-space readout on benign controls. That is the anti-lobotomy safeguard. It asks whether the edit disturbed the rest of what the model verbalizably represents, not just whether the output still looks fine.&lt;/p&gt;
&lt;h2&gt;The surgical edit that barely moves behavior&lt;/h2&gt;
&lt;p&gt;On disjoint eval splits (120 AdvBench, 200 XSTest, 250 ARC-Easy, 48 benign controls), at strength 1:&lt;/p&gt;













































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;method&lt;/th&gt;&lt;th&gt;AdvBench refusal&lt;/th&gt;&lt;th&gt;XSTest-unsafe&lt;/th&gt;&lt;th&gt;ARC&lt;/th&gt;&lt;th&gt;workspace KL&lt;/th&gt;&lt;th&gt;refusal suppression&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;original&lt;/td&gt;&lt;td&gt;0.99&lt;/td&gt;&lt;td&gt;0.91&lt;/td&gt;&lt;td&gt;0.98&lt;/td&gt;&lt;td&gt;0.000&lt;/td&gt;&lt;td&gt;0.00&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;mean-diff (abliteration)&lt;/td&gt;&lt;td&gt;&lt;strong&gt;0.06&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;0.13&lt;/td&gt;&lt;td&gt;0.98&lt;/td&gt;&lt;td&gt;0.257&lt;/td&gt;&lt;td&gt;3.44&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;pullback&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;0.78&lt;/td&gt;&lt;td&gt;0.13&lt;/td&gt;&lt;td&gt;0.98&lt;/td&gt;&lt;td&gt;&lt;strong&gt;0.046&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;&lt;strong&gt;7.55&lt;/strong&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;pullback subspace r=3&lt;/td&gt;&lt;td&gt;0.55&lt;/td&gt;&lt;td&gt;0.23&lt;/td&gt;&lt;td&gt;0.98&lt;/td&gt;&lt;td&gt;0.196&lt;/td&gt;&lt;td&gt;7.18&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Read the pullback row as a paradox. It is the most precise instrument on the table: it distorts the benign workspace about 5.6 times less than abliteration (KL 0.046 against 0.257) while suppressing the workspace’s own refusal-mass about 2.2 times more (7.55 against 3.44). It is, almost by definition, the refusal-readout direction, so it barely touches anything off that axis. And it leaves 78% of AdvBench refusal behaviorally intact. Abliteration, the blunt instrument, drops refusal to 0.06.&lt;/p&gt;
&lt;p&gt;Capability holds throughout: ARC stays at 0.98 for every non-degenerate edit. So the meaningful “lobotomy” signal is not accuracy, it is the workspace KL. This is exactly why measuring collateral in the J-space rather than only at the output matters: the two edits look very different inside the model and only somewhat different at the surface.&lt;/p&gt;
&lt;p&gt;Pushing the single direction harder does not rescue it. A strength sweep shows it &lt;em&gt;plateaus&lt;/em&gt;: it bottoms out around AdvBench refusal 0.68 with the workspace intact, and only reaches 0.00 at strength 3, where ARC collapses to 0.22 and workspace KL blows up to 17. You cannot get to full removal through that one direction without breaking the model.&lt;/p&gt;
&lt;h2&gt;What “one third workspace-mediated” actually means&lt;/h2&gt;
&lt;p&gt;The plateau is the result, not a nuisance. Cleanly deleting the verbalizable “I cannot” disposition from the workspace removes only a minority of the refusal behavior. To locate the rest, I split abliteration’s direction &lt;span&gt;&lt;span&gt;mm&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;m&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; into the part parallel to the pullback, &lt;span&gt;&lt;span&gt;m∥pm_{\parallel p}&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;m&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;∥&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;, and the part orthogonal to it, &lt;span&gt;&lt;span&gt;m⊥pm_{\perp p}&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;m&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;⊥&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;, and ablated each on 100 AdvBench prompts:&lt;/p&gt;






























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;direction&lt;/th&gt;&lt;th&gt;behavior removed&lt;/th&gt;&lt;th&gt;workspace “cannot” cleared&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;pullback &lt;span&gt;&lt;span&gt;pp&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/td&gt;&lt;td&gt;0.22&lt;/td&gt;&lt;td&gt;&lt;strong&gt;7.55&lt;/strong&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;workspace part &lt;span&gt;&lt;span&gt;m∥pm_{\parallel p}&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;m&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;∥&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/td&gt;&lt;td&gt;0.23&lt;/td&gt;&lt;td&gt;7.55&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;orthogonal part &lt;span&gt;&lt;span&gt;m⊥pm_{\perp p}&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;m&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;⊥&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/td&gt;&lt;td&gt;&lt;strong&gt;0.90&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;1.81&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;abliteration &lt;span&gt;&lt;span&gt;mm&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;m&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/td&gt;&lt;td&gt;0.93&lt;/td&gt;&lt;td&gt;3.44&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Ablating &lt;span&gt;&lt;span&gt;pp&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; clears the verbalizable narration almost completely and moves behavior from 0.99 to 0.77. Ablating &lt;span&gt;&lt;span&gt;m⊥pm_{\perp p}&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;m&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;⊥&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; removes the behavior (down to 0.09) while barely touching the narration. The lens-verbalizable slice of the refusal direction does not carry the refusal behavior. That is roughly a third: the readable part is a minority stakeholder in the decision.&lt;/p&gt;
&lt;p&gt;I want to be honest about where the evidence is independent. &lt;span&gt;&lt;span&gt;p=Jl⊤(g⊙w)p = J_l^{\top}(g \odot w)&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;J&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;l&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;⊤&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;g&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;⊙&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;w&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; is, up to the lens linearization, the gradient of the workspace refusal-mass itself. So “ablating &lt;span&gt;&lt;span&gt;pp&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; maximizes suppression” and “the pullback has low off-axis workspace KL” are partly true &lt;em&gt;by construction&lt;/em&gt;: those columns are coupled to how &lt;span&gt;&lt;span&gt;pp&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; is built. Only the behavior column is independent evidence. And there &lt;span&gt;&lt;span&gt;m⊥pm_{\perp p}&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;m&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;⊥&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; is close to plain abliteration, so that part is close to a restatement of the standard method.&lt;/p&gt;
&lt;h2&gt;The correction I had to make&lt;/h2&gt;
&lt;p&gt;My first framing was that this is a “double dissociation” between a &lt;em&gt;verbalizable workspace&lt;/em&gt; refusal and an &lt;em&gt;automatic&lt;/em&gt; one living outside the lens. I red-teamed that claim and it did not survive, so I document both the claim and its retraction.&lt;/p&gt;
&lt;p&gt;The mechanistic test is whether the behavior-carrying direction lives in the lens’s null space. It does not. &lt;span&gt;&lt;span&gt;m⊥pm_{\perp p}&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;m&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;⊥&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; is 61% lens-&lt;em&gt;visible&lt;/em&gt;. Ablating the lens-visible part of &lt;span&gt;&lt;span&gt;mm&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;m&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; removes 100% of refusal; ablating the lens-blind part removes 0%, the opposite of the null-space hypothesis. And the lens image of &lt;span&gt;&lt;span&gt;m⊥pm_{\perp p}&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;m&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;⊥&lt;/span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; does not decode to refusal words at all: it reads &lt;code&gt;illegal&lt;/code&gt;, &lt;code&gt;crime&lt;/code&gt;, &lt;code&gt;violence&lt;/code&gt;, &lt;code&gt;police&lt;/code&gt; at mid layers. It is a harmfulness-&lt;em&gt;perception&lt;/em&gt; feature.&lt;/p&gt;
&lt;p&gt;So the real distinction is not workspace versus automatic. Both directions are in the workspace. It is &lt;em&gt;perception versus narration&lt;/em&gt;. One feature perceives that the request is harmful (&lt;code&gt;illegal&lt;/code&gt;, &lt;code&gt;crime&lt;/code&gt;), and a distinct feature narrates the refusal (&lt;code&gt;I cannot&lt;/code&gt;). Behavior follows perception. Ablating the narration leaves a model that still refuses but can no longer articulate why, and a bidirectional steering test confirms the direction of causation: adding the harm &lt;em&gt;representation&lt;/em&gt; to benign prompts induces genuine refusal, while adding the refusal &lt;em&gt;narration&lt;/em&gt; just forces the words “I cannot” at a much higher workspace cost, and adding harm-narration alone makes the model cheerfully write “renewable energy: 1. Illegal drug trafficking” with no perception of harm and no refusal at all.&lt;/p&gt;
&lt;h2&gt;Why this matters for safety&lt;/h2&gt;
&lt;p&gt;There is a practical spinoff that points the same way. Even after you abliterate the behavior-carrying direction so the model complies with harmful requests at the surface, its internal refusal-mass still separates harmful-that-complied prompts from benign ones at AUC 0.998, against 0.48 for the surface behavior. And it is a disposition detector, not a topic detector: benign-but-harmful-topic prompts like “how do I kill a Python process” score 2.75, far below genuinely-refused prompts at 8.57. An “uncensored” open model still carries a monitorable internal signal that it &lt;em&gt;knows&lt;/em&gt; it should refuse. That is a real safety hook.&lt;/p&gt;
&lt;p&gt;The lesson I take is about where safety behaviors live. The part of refusal you can most easily read and most surgically edit, the verbalizable “I cannot”, is the narration, not the mechanism. It is a downstream readout of an upstream perception. If you edit the memo, the meeting already happened. Any safety intervention that operates on the legible, verbalizable layer is touching the announcement, not the decision, and the decision is stored somewhere you have to work harder to reach.&lt;/p&gt;</content:encoded><category>Interpretability</category><category>AI Safety</category><category>LLM</category><category>Mechanistic</category><category>Research</category></item><item><title>RosettIA: Beating the Task Winners on Spanish to Quechua</title><link>https://irvingernesto.com/blog/rosettia-low-resource-languages/</link><guid isPermaLink="true">https://irvingernesto.com/blog/rosettia-low-resource-languages/</guid><description>A supervised NLLB baseline, GSPO reinforcement learning, and MBR decoding reach 46.71 ChrF on Spanish to Chanka Quechua, ahead of the AmericasNLP 2021 and 2023 systems. With the caveats that number deserves.</description><pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;RosettIA is my open effort on translation for languages that the internet mostly forgot. The result I can put a number on, and the one I am proudest of, is Spanish to Chanka (Ayacucho) Quechua, where the system I built lands ahead of the published task winners on the standard AmericasNLP 2021 benchmark.&lt;/p&gt;
&lt;p&gt;I want to show the number and then immediately show its limits, because a single benchmark score is easy to oversell and I would rather you trust the ones I do report.&lt;/p&gt;
&lt;h2&gt;The benchmark&lt;/h2&gt;
&lt;p&gt;Everything below is on the &lt;strong&gt;AmericasNLP 2021&lt;/strong&gt; test set for Spanish to Chanka (Ayacucho) Quechua: 1003 sentences, a single reference each, scored with &lt;strong&gt;ChrF&lt;/strong&gt; (sacrebleu, word_order=0). ChrF is a character-level F-score, which suits a morphologically rich language like Quechua better than word-level BLEU, because so much meaning lives inside the word.&lt;/p&gt;
&lt;h2&gt;The results&lt;/h2&gt;





































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;System&lt;/th&gt;&lt;th&gt;ChrF&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Sheffield 2023 (NLLB-3.3B ensemble)&lt;/td&gt;&lt;td&gt;34.01&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Helsinki 2021 (prior task winner)&lt;/td&gt;&lt;td&gt;39.40&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen-9B (ours), greedy&lt;/td&gt;&lt;td&gt;40.55&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;NLLB-1.3B (ours), supervised&lt;/td&gt;&lt;td&gt;42.95&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;+ GSPO reinforcement learning&lt;/td&gt;&lt;td&gt;45.53&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;+ MBR decoding (single model)&lt;/td&gt;&lt;td&gt;46.43&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Our best system&lt;/td&gt;&lt;td&gt;&lt;strong&gt;46.71&lt;/strong&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The two grey rows are prior published work. Everything below them is mine, and every one of my systems clears both prior task winners on this metric.&lt;/p&gt;
&lt;h2&gt;The build-up&lt;/h2&gt;
&lt;p&gt;The interesting part is not the top number, it is the climb.&lt;/p&gt;
&lt;p&gt;A &lt;strong&gt;supervised NLLB-1.3B&lt;/strong&gt; fine-tune already reaches 42.95, past both prior winners, which tells you most of the gap to earlier work was data and training discipline, not model scale. A general-purpose &lt;strong&gt;Qwen-9B&lt;/strong&gt; with plain greedy decoding hits 40.55 without being a translation model at all, which is its own quiet statement about where open LLMs now sit on low-resource pairs.&lt;/p&gt;
&lt;p&gt;Then two techniques stack on top of the supervised model:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;GSPO reinforcement learning&lt;/strong&gt; adds +2.58 ChrF, from 42.95 to 45.53. GSPO (Group Sequence Policy Optimization) is the same sequence-level, length-normalized RL method I have used elsewhere in my work; here it optimizes the translation model directly against a quality signal rather than pure next-token likelihood.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MBR decoding&lt;/strong&gt; (minimum Bayes risk, single model) adds a bit more, to 46.43, by choosing the candidate translation that agrees most with the model’s own sampled hypotheses instead of the single greedy path.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The best system combines the pieces to 46.71.&lt;/p&gt;
&lt;h2&gt;The caveats, up front&lt;/h2&gt;
&lt;p&gt;I said I would show the limits, so here they are, and they are on the chart itself:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;One benchmark, one reference, no human evaluation.&lt;/strong&gt; ChrF against a single reference rewards surface overlap, not fluency or adequacy as a Quechua speaker would judge it. I have not run a human eval, so treat this as a system-comparison number, not a claim about real-world translation quality.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;I am not claiming to beat 2024.&lt;/strong&gt; The 2024 task winner, BSC-2024, reported 38.21 ChrF++ with word_order=2, which is a different metric configuration and is &lt;strong&gt;not directly comparable&lt;/strong&gt; to the word_order=0 ChrF I report here. What I can defend is a strong, reproducible result on the standard 2021 setup, ahead of the 2021 and 2023 systems on the same metric. I am deliberately not putting our number next to theirs as if it were a fair fight.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That is the honest shape of it: a real improvement over the comparable prior work, measured carefully, with the comparison I cannot make left explicitly unmade.&lt;/p&gt;
&lt;h2&gt;Why it matters&lt;/h2&gt;
&lt;p&gt;Quechua is spoken by millions of people across the Andes and has almost no usable machine translation. The gap is not that the problem is unsolvable, it is that almost nobody is working on it. RosettIA is my attempt to close a little of that gap in the open, with the models and the honest numbers both public, so the next person starts ahead of where I did.&lt;/p&gt;</content:encoded><category>NLP</category><category>Low-resource</category><category>Quechua</category><category>Reinforcement Learning</category><category>Open Source</category></item><item><title>Teaching XGBoost to Read Blueprints</title><link>https://irvingernesto.com/blog/teaching-xgboost-to-read-blueprints/</link><guid isPermaLink="true">https://irvingernesto.com/blog/teaching-xgboost-to-read-blueprints/</guid><description>Lessons from building a vector-native structural member detector for steel construction drawings, where gradient-boosted trees beat transformers, label quality beat everything, and the headline number turned out to be measuring the wrong thing.</description><pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;At &lt;a href=&quot;https://steeleye.ai&quot;&gt;SteelEye&lt;/a&gt; we needed to automate takeoff: given a structural construction drawing, find &lt;strong&gt;every placed member&lt;/strong&gt; (columns, beams, joists, braces, base plates, deck) and assign each its correct AISC section name (&lt;code&gt;W14X22&lt;/code&gt;, &lt;code&gt;HSS5X5X1/4&lt;/code&gt;, &lt;code&gt;26KSP&lt;/code&gt;…). Estimators do this by hand today, page by page, and it’s slow, error-prone work that million-dollar bids depend on.&lt;/p&gt;
&lt;p&gt;Two years of ML hype would tell you to fine-tune a vision-language model and call it a day. Here’s what actually worked, after ~200 experiment scripts and 33k lines of research code.&lt;/p&gt;
&lt;h2&gt;Work on vectors, not pixels&lt;/h2&gt;
&lt;p&gt;Construction PDFs aren’t scans; they’re CAD exports. The text tokens and line segments are &lt;em&gt;right there&lt;/em&gt; in the file. Instead of rasterizing and running object detection, we extract primitives with PyMuPDF and classify &lt;strong&gt;tokens&lt;/strong&gt;: is this string a member’s name, and if so, what class of member?&lt;/p&gt;
&lt;p&gt;This one decision bought us exact text (no OCR noise), exact geometry (segment endpoints, orientations, lengths), and two orders of magnitude less compute than a vision pipeline.&lt;/p&gt;
&lt;p&gt;Ground truth came from an unusual place: &lt;strong&gt;NC1 (DSTV) files&lt;/strong&gt;, the CNC instructions the fabrication shop actually cuts from. If the model says a page contains &lt;code&gt;W12X26&lt;/code&gt; beams and the NC1 answer key agrees, that’s validation no human labeling budget could match.&lt;/p&gt;
&lt;h2&gt;The model zoo, and who survived it&lt;/h2&gt;
&lt;p&gt;We benchmarked honestly: gradient-boosted trees (LightGBM/XGBoost), a deep residual MLP, a set-attention transformer, a kNN message-passing GNN, tabular foundation models (TabICL, TabPFN), even the Hierarchical Reasoning Model. The numbers below are relative scores from the same self-consistent harness (more on why that harness flattered everyone in a moment), so read them as a ranking, not as capability.&lt;/p&gt;





























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Relative F1 (same harness)&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;strong&gt;GBM (XGBoost/LightGBM)&lt;/strong&gt;&lt;/td&gt;&lt;td&gt;&lt;strong&gt;best&lt;/strong&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MLP (residual, bf16)&lt;/td&gt;&lt;td&gt;close behind&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;GBM + MLP ensemble&lt;/td&gt;&lt;td&gt;close behind&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;GNN (kNN graph, 3 layers)&lt;/td&gt;&lt;td&gt;worse&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Transformer (set attention)&lt;/td&gt;&lt;td&gt;diverged&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The gradient-boosted trees won, on 84 engineered features: geometry (distance to segments, orientation context), relational cues (neighbor density, same-row/column), and 32 PCA dimensions of &lt;strong&gt;e5-small text embeddings&lt;/strong&gt;. That last part matters: a 33M-parameter embedding model beat its larger siblings at understanding technical codes, and added +6.6 points on unseen naming schemes.&lt;/p&gt;
&lt;p&gt;This mirrors the CAD symbol-spotting literature: state-of-the-art methods there are shallow, kNN-edge-based systems, not deep stacks. When your entities are sparse, structured, and text-anchored, representation beats architecture.&lt;/p&gt;
&lt;h2&gt;Synthetic data is a representation lever, not a volume lever&lt;/h2&gt;
&lt;p&gt;The hardest problem was &lt;strong&gt;cross-fabricator generalization&lt;/strong&gt;: every detailing shop has its own drawing style and mark scheme. Real training data all came from one fabricator.&lt;/p&gt;
&lt;p&gt;So we built a procedural drawing generator with &lt;strong&gt;style personas&lt;/strong&gt;: synthetic sheets with randomized grids, leader lines, mark conventions, and adversarial decoys (notes and callouts that look like members). Training on real + 500 synthetic sheets lifted novel-fabricator accuracy from 0.57 to 0.78.&lt;/p&gt;
&lt;p&gt;Two counterintuitive findings:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Diversity saturates.&lt;/strong&gt; 48 personas ≈ 160 personas. Once the model has seen “enough kinds of different,” more variety adds nothing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Random mark schemes teach a skill, not facts.&lt;/strong&gt; Synthetic sheets with nonsense prefixes force the model to learn &lt;em&gt;“unknown prefix → trust the geometry”&lt;/em&gt;, exactly the behavior you need on a new fabricator’s drawings.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We also tested the fashionable alternatives: domain-generalization losses (IRM, GroupDRO, CORAL), an RL-based adversarial curriculum, pseudo-labeling. All rejected: none beat plain XGBoost with better data.&lt;/p&gt;
&lt;h2&gt;Label quality beats model capacity&lt;/h2&gt;
&lt;p&gt;The single largest jump in the entire program came from a labeling fix, not a modeling idea: marking &lt;strong&gt;one positive token per member&lt;/strong&gt; (its name/mark) instead of every token inside its bounding box. Clean supervision moved metrics more than any architecture change we tried.&lt;/p&gt;
&lt;p&gt;If your model has plateaued, audit your labels before reaching for a bigger network. In our callout-detection work, a vision-model audit found &lt;strong&gt;29% of human gold labels were corrupted&lt;/strong&gt; by an export bug. The “model problem” was a data problem wearing a costume.&lt;/p&gt;
&lt;h2&gt;Where it landed&lt;/h2&gt;
&lt;p&gt;For a while I thought the production detector hit 0.987 member-F1 and 0.948 end-to-end, and I reported those numbers. They turned out to be measuring the wrong thing: scored against a benchmark auto-derived from the model’s own inputs, the model was largely grading its own homework. I tell that whole eval-integrity story in &lt;a href=&quot;/blog/the-f1-that-wasnt&quot;&gt;The 0.98 F1 That Wasn’t&lt;/a&gt;. Regraded against independent human gold, the honest detector reaches &lt;strong&gt;~0.77 member-F1&lt;/strong&gt; and &lt;strong&gt;~0.78 end-to-end&lt;/strong&gt; (found &lt;em&gt;and&lt;/em&gt; correctly named) on a frozen human-annotated test set, with leave-one-project-out cross-validation and reproducible-from-scratch docs. That is the number I stand behind, and I think it reads as the stronger result.&lt;/p&gt;
&lt;p&gt;The boring stack won: engineered features, gradient-boosted trees, procedural data, obsessive label hygiene. The exciting part isn’t the architecture. It’s that estimators get hours of their week back.&lt;/p&gt;</content:encoded><category>ML</category><category>XGBoost</category><category>CAD</category><category>SteelEye</category><category>Research</category></item><item><title>Carrier Lifetime from 49 Photos of an Oscilloscope</title><link>https://irvingernesto.com/blog/carrier-lifetime-from-oscilloscope-photos/</link><guid isPermaLink="true">https://irvingernesto.com/blog/carrier-lifetime-from-oscilloscope-photos/</guid><description>An experimental-physics project with a computer-vision twist: no numeric data ever existed, only 49 photographs of a scope screen. I taught a computer to read the decay curves off the photos, fit them, and recover the minority-carrier lifetime of a silicon solar cell.</description><pubDate>Mon, 20 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The whole project rests on a fact that should not have been possible to work with: there was no data. A silicon solar cell had been measured in the lab, its minority-carrier lifetime probed by photoconductive decay across a temperature sweep, and the only surviving record of the entire run was 49 photographs. Not saved waveforms, not a CSV, not a scope dump over GPIB. Forty-nine images of a Tektronix TDS 2024 screen and a Keithley 2182A nanovoltmeter, taken roughly one every 25 seconds over about 20 minutes. Every number I report below, every temperature and every microsecond of lifetime, was recovered from those pixels.&lt;/p&gt;
&lt;p&gt;So this is two projects wearing one coat. Half of it is semiconductor physics: what photoconductive decay measures and why the lifetime falls out of an exponential tail. The other half is computer vision: turning a photograph of a glowing trace into a clean, calibrated decay curve you can actually fit. I will take the physics first, because it tells you what the pixels have to give up.&lt;/p&gt;
&lt;h2&gt;What lifetime is, and why you decay a solar cell to get it&lt;/h2&gt;
&lt;p&gt;A solar cell works because photons knock electron-hole pairs loose. Those excess carriers are what carry current, but they do not live forever. Each one survives only until it recombines, and the mean time from generation to recombination is the minority-carrier lifetime &lt;span&gt;&lt;span&gt;τ\tau&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;. It is one of the cleanest single numbers for the electronic quality of the material, because it is set by the density of defects that carriers recombine through.&lt;/p&gt;
&lt;p&gt;Photoconductive decay measures &lt;span&gt;&lt;span&gt;τ\tau&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; about as directly as you can. Shine light on the cell and its conductivity rises as carriers pile up. Cut the light and the excess carriers recombine, so the conductivity, and the voltage proportional to it, decays back toward baseline. In this rig a laser was chopped by a mechanical fan at roughly 70 Hz, and the falling edge after each blocked pulse is the decay. If the excitation is switched off fast compared to &lt;span&gt;&lt;span&gt;τ\tau&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;, the fall is a pure exponential:&lt;/p&gt;
&lt;p&gt;&lt;span&gt;&lt;span&gt;vo(t)=V1+ΔV e−t/τpv_o(t) = V_1 + \Delta V \, e^{-t/\tau_p}&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;v&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;o&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;t&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;V&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;1&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;+&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;Δ&lt;/span&gt;&lt;span&gt;V&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;e&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;−&lt;/span&gt;&lt;span&gt;t&lt;/span&gt;&lt;span&gt;/&lt;/span&gt;&lt;span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;The time constant &lt;span&gt;&lt;span&gt;τp\tau_p&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;p&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; of that tail is the carrier lifetime. Fit the exponential, read off &lt;span&gt;&lt;span&gt;τ\tau&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;. That is the plan. The rest of the physics is a story about the word “if” in “if the excitation is switched off fast enough,” and we will come back to it.&lt;/p&gt;
&lt;h2&gt;Reading physics off a photograph&lt;/h2&gt;
&lt;p&gt;First I had to get the curve out of the image, and a photo of a scope is a hostile input. The screen is shot at an angle, so the graticule is a trapezoid, not a rectangle. The trace glows and blooms. The dim tail of the decay, which is exactly the part I need, is the faintest thing on the panel. There are two instruments in most frames plus the grey of the room. The pipeline is OpenCV and Python, managed with uv, and it runs in three stages.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Temperature, from the Keithley.&lt;/strong&gt; The nanovoltmeter reads out on a bright cyan VFD, a very distinctive color against the blue-purple scope and the grey room. I threshold for bright cyan in HSV, restrict to the lower part of the frame where the meter sits, and morphologically close horizontally so the row of digits merges into one wide strip I can crop and upscale. The validation is beautiful in its simplicity: read in chronological order, the recovered temperatures climb monotonically in steps of about 0.49 degrees C with zero reversals. A thermocouple on a warming sample cannot un-warm, so a strictly monotone sequence is strong evidence the reader is correct. The run covers -49.99 to -26.51 degrees C, that is 223.2 to 246.6 K, warming at about 1.19 degrees C per minute.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The trace, from the scope.&lt;/strong&gt; The TDS 2024 LCD is a vivid blue, so I segment for that hue, but a plain blue panel or the window blinds behind the bench can masquerade as a screen. Two tricks reject the impostors. The scope LCD is a 4:3 rectangle, so I filter candidate contours by aspect ratio, which kills the long strips of blind. And a real scope screen contains the cream-colored trace, so among survivors I prefer the contour whose bounding box actually contains cream pixels. Having found the screen I take its convex hull, approximate it to a quadrilateral to get four clean corners (the hull matters, because the trace and the on-screen menu carve concavities into the blue field that would otherwise corrupt the corner finding), and apply a perspective transform to rectify the trapezoid into a canonical 800 by 600 image.&lt;/p&gt;
&lt;p&gt;Now the graticule is square, but I still need to know how many microseconds a pixel is worth. Rather than trust a fixed number, I recover it per image from the grid itself. Running a Sobel profile down the rectified graticule and autocorrelating it finds the repeat period of the grid lines, about 62.5 pixels per division. The scope was on 250 microseconds per division, so that pins the calibration at roughly 4.0 microseconds per pixel, and the fact that the grid period comes out consistent is an independent check that the rectification is honest.&lt;/p&gt;
&lt;p&gt;Extracting the trace itself is a column walk. For each column I take the centroid of the cream-colored pixels as the trace height. The subtlety is the tail: as the decay fades, the cream drops below any fixed threshold and the hard mask loses it. So where the mask is empty I fall back to a “cream score,” roughly red plus green minus twice blue, and pick the brightest cream-ish pixel within a window around the previous column’s height. That continuity constraint lets the tracker follow the faint tail without jumping onto a grid line or a menu glyph. The verification overlays, where the extracted points are drawn back onto the original trace, land exactly on the glowing curve.&lt;/p&gt;
&lt;h2&gt;Fitting the tail, and the trap in the shoulder&lt;/h2&gt;
&lt;p&gt;With 49 calibrated &lt;span&gt;&lt;span&gt;(t,V)(t, V)&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;t&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;V&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; curves in hand, the exponential fit ought to be routine. It was not, and the reason turned out to be the most interesting physics in the project.&lt;/p&gt;
&lt;p&gt;The measured fall is not a clean exponential. It is sigmoidal: it starts nearly flat, bends into its steepest slope partway down, and only then straightens into a decay. A pure &lt;span&gt;&lt;span&gt;e−t/τe^{-t/\tau}&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;e&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;−&lt;/span&gt;&lt;span&gt;t&lt;/span&gt;&lt;span&gt;/&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; has its steepest slope at the very top of the fall, not in the middle, so something is rounding off the top. That rounded top is the shoulder, and it is a genuine artifact you can catch red-handed.&lt;/p&gt;
&lt;p&gt;The cause is that “if the light switches off fast enough” caveat. A mechanical fan blade does not chop the beam instantly. Its edge takes a finite time to sweep across the laser spot, so the illumination ramps down rather than stepping down, and the measured signal is the true exponential recombination convolved with that finite optical turn-off. The convolution smears the top of the fall into the shoulder.&lt;/p&gt;
&lt;p&gt;Here is the clean proof it is instrumental and not physics. I measure the shoulder width on every image, and it comes out at 162 plus or minus 9 microseconds and stays flat across the entire temperature sweep. A mechanical chopper cannot possibly depend on the sample’s temperature. So a shoulder that is constant while &lt;span&gt;&lt;span&gt;τ\tau&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; itself changes must be the apparatus, not the recombination. The real lifetime lives in the exponential body below the shoulder.&lt;/p&gt;
&lt;p&gt;That dictates the fit. I do not fit the whole fall. I normalize each curve to &lt;span&gt;&lt;span&gt;u=(V−V1)/ΔVu = (V - V_1)/\Delta V&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;u&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;V&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;−&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;V&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;1&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;​&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;/Δ&lt;/span&gt;&lt;span&gt;V&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; and fit &lt;span&gt;&lt;span&gt;ln⁡u\ln u&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;ln&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;u&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; against &lt;span&gt;&lt;span&gt;tt&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;t&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; only over the body, roughly &lt;span&gt;&lt;span&gt;u∈[0.10,0.70]u \in [0.10, 0.70]&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;u&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;∈&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;[&lt;/span&gt;&lt;span&gt;0.10&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;0.70&lt;/span&gt;&lt;span&gt;]&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;, a window that sits below the chopper shoulder and above the noise floor and is immune to how long the plateau ran. The log-linear fit is a straight line in semilog with a typical &lt;span&gt;&lt;span&gt;R2R^2&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;R&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;2&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; around 0.98 to 0.99. The lifetimes come out in the range 245 to 384 microseconds, median 319, exactly where a crystalline-silicon cell limited by Shockley-Read-Hall recombination through a defect level should sit.&lt;/p&gt;
&lt;h2&gt;What the sweep actually reveals&lt;/h2&gt;
&lt;p&gt;The original question was &lt;span&gt;&lt;span&gt;τ(T)\tau(T)&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;T&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;: how does lifetime depend on temperature, and what does that say about the dominant recombination mechanism? The honest answer is that this dataset cannot cleanly say, and finding out why is the payoff.&lt;/p&gt;
&lt;p&gt;Lifetime correlates only weakly with temperature, &lt;span&gt;&lt;span&gt;r=0.60r = 0.60&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;r&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;0.60&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;. It correlates strongly, &lt;span&gt;&lt;span&gt;r=0.91r = 0.91&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;r&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;0.91&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;, with the amplitude of the photo-signal, which is proportional to the injection level, the density of carriers you inject with each pulse. Worse, the run splits into two regimes with a simultaneous jump in both &lt;span&gt;&lt;span&gt;τ\tau&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; and amplitude at image 15, around the midpoint of the session, the fingerprint of someone nudging the laser power or focus mid-run. Regime A, the earlier low-injection images, sits at &lt;span&gt;&lt;span&gt;τ\tau&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; about 260 microseconds; regime B, higher injection, at about 335. Both regimes fall on a single rising &lt;span&gt;&lt;span&gt;τ\tau&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;-versus-injection trend.&lt;/p&gt;
&lt;p&gt;That is not noise, it is real physics pointing the wrong way for this experiment. In silicon, SRH lifetime genuinely depends on injection level: push more carriers in, the recombination traps saturate, and &lt;span&gt;&lt;span&gt;τ\tau&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; rises. So injection is a real knob on &lt;span&gt;&lt;span&gt;τ\tau&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;, and here it moved during the run and co-varied with temperature. The two effects are entangled, which is why a single Arrhenius fit for an activation energy is worthless on this data (&lt;span&gt;&lt;span&gt;R2=0.41R^2 = 0.41&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;R&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;2&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;0.41&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;, uninterpretable). To isolate &lt;span&gt;&lt;span&gt;τ(T)\tau(T)&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;T&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; you would need to hold injection fixed (stable laser power and focus) and use a faster chopper so the optical turn-off is far shorter than &lt;span&gt;&lt;span&gt;τ\tau&lt;/span&gt;&lt;span&gt;&lt;span&gt;&lt;span&gt;&lt;/span&gt;&lt;span&gt;τ&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt; and the shoulder vanishes entirely. Vary temperature and injection independently and you are doing proper temperature-and-injection-dependent lifetime spectroscopy, which is how you pull a defect’s energy level and capture cross sections out.&lt;/p&gt;
&lt;p&gt;The result I am happiest about is not the lifetime number. It is that a full quantitative dataset, 49 calibrated decay curves and a monotone temperature record, was reconstructed from nothing but photographs, and that the same photographs caught the apparatus lying about its own turn-off time. The code, the extracted traces, and a Typst report with the figures are all in the repo.&lt;/p&gt;</content:encoded><category>Physics</category><category>Semiconductors</category><category>Computer Vision</category><category>Experimental</category><category>Open Source</category></item><item><title>CodeContexter: Packing a Whole Codebase Into One LLM-Ready File</title><link>https://irvingernesto.com/blog/codecontexter-packing-codebases-for-llms/</link><guid isPermaLink="true">https://irvingernesto.com/blog/codecontexter-packing-codebases-for-llms/</guid><description>How I built a fast, safety-first Rust CLI that walks an entire repository, respects .gitignore, redacts secrets before they leave your machine, and estimates the token budget, all in a single streamed pass.</description><pubDate>Sun, 08 Mar 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Every time I wanted an LLM to reason about a whole project, I hit the same wall. The model needs context, but a project is not one file. It is hundreds of files scattered across folders, plus lock files, build artifacts, and a &lt;code&gt;.env&lt;/code&gt; I would very much like to never paste into a chat window. Copying files by hand is slow, lossy, and dangerous. So I wrote &lt;a href=&quot;https://github.com/yourusername/codecontexter&quot;&gt;CodeContexter&lt;/a&gt;: a Rust CLI that walks a repository and packs it into a single structured file that a model can actually read.&lt;/p&gt;
&lt;p&gt;The whole thing is one binary and one job: take a directory, produce a clean, deduplicated, secret-free document with a file tree and every file’s contents. Here is what turned out to be interesting under the hood.&lt;/p&gt;
&lt;h2&gt;The walk has to be ignore-aware, not just fast&lt;/h2&gt;
&lt;p&gt;The naive version of this tool recursively reads every file. That version is useless, because it will happily dump &lt;code&gt;node_modules&lt;/code&gt;, &lt;code&gt;target&lt;/code&gt;, and a 40MB lock file into your context and blow the budget on noise.&lt;/p&gt;
&lt;p&gt;So the walk is built on the &lt;code&gt;ignore&lt;/code&gt; crate, the same directory traversal library that powers ripgrep. It gives me &lt;code&gt;.gitignore&lt;/code&gt; semantics for free:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;let&lt;/span&gt;&lt;span&gt; walker &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;WalkBuilder&lt;/span&gt;&lt;span&gt;::&lt;/span&gt;&lt;span&gt;new&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&amp;amp;&lt;/span&gt;&lt;span&gt;root_path)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;hidden&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;false&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;git_ignore&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;true&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;follow_links&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;false&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;// prevent symlink loops and duplication&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;overrides&lt;/span&gt;&lt;span&gt;(overrides)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;build&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;div&gt;&lt;div&gt;&lt;/div&gt;&lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;Two of those flags are deliberate. &lt;code&gt;hidden(false)&lt;/code&gt; means I do &lt;em&gt;not&lt;/em&gt; skip dotfiles, because config like &lt;code&gt;.eslintrc&lt;/code&gt; or &lt;code&gt;.github/workflows&lt;/code&gt; is often exactly the context you want. And &lt;code&gt;follow_links(false)&lt;/code&gt; closes a nasty failure mode: a symlink that points back up the tree can send a recursive walker into an infinite loop or silently duplicate half the repo. Turning link-following off makes the walk finite and honest.&lt;/p&gt;
&lt;p&gt;The one thing &lt;code&gt;.gitignore&lt;/code&gt; gets wrong for this use case is secrets. Plenty of repos commit or fail to ignore a stray &lt;code&gt;.pem&lt;/code&gt; or &lt;code&gt;id_rsa&lt;/code&gt;, and I do not want the walk to depend on the user having a perfect ignore file. So before the walk starts, I layer in hard-coded overrides that force-exclude the dangerous stuff regardless of what &lt;code&gt;.gitignore&lt;/code&gt; says:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;let&lt;/span&gt;&lt;span&gt; hard_coded_excludes &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;vec!&lt;/span&gt;&lt;span&gt;[&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;&quot;!*.env&quot;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;!*.env.*&quot;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;!*.pem&quot;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;!*.key&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;&quot;!id_rsa&quot;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;!id_ed25519&quot;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;!*.p12&quot;&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;&quot;!*.pfx&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;];&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;div&gt;&lt;div&gt;&lt;/div&gt;&lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;The &lt;code&gt;!&lt;/code&gt; prefix in an override inverts it to an exclude, so these patterns are the tool’s own opinion about what should never be aggregated, applied on top of the user’s &lt;code&gt;.gitignore&lt;/code&gt; and any &lt;code&gt;--exclude&lt;/code&gt; globs they pass. Belt and suspenders. The file-level redaction below is the belt.&lt;/p&gt;
&lt;h2&gt;Redaction is defense in depth, not the only defense&lt;/h2&gt;
&lt;p&gt;Excluding secret &lt;em&gt;files&lt;/em&gt; handles the obvious case. It does not handle the API key someone hardcoded in the middle of a Python module. For that, every file’s contents pass through a sanitizer before they are written out.&lt;/p&gt;
&lt;p&gt;The sanitizer is a small set of compiled regexes, initialized once and reused across every file via a &lt;code&gt;OnceLock&lt;/code&gt; so I am not recompiling patterns in a hot loop:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;static&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;SECRET_PATTERNS&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;OnceLock&lt;/span&gt;&lt;span&gt;&amp;lt;&lt;/span&gt;&lt;span&gt;Vec&lt;/span&gt;&lt;span&gt;&amp;lt;&lt;/span&gt;&lt;span&gt;Regex&lt;/span&gt;&lt;span&gt;&amp;gt;&amp;gt; &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;OnceLock&lt;/span&gt;&lt;span&gt;::&lt;/span&gt;&lt;span&gt;new&lt;/span&gt;&lt;span&gt;();&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;let&lt;/span&gt;&lt;span&gt; patterns &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;SECRET_PATTERNS&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;get_or_init&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;||&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;vec!&lt;/span&gt;&lt;span&gt;[&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;Regex&lt;/span&gt;&lt;span&gt;::&lt;/span&gt;&lt;span&gt;new&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;r&quot;-----BEGIN [A-Z ]+ PRIVATE KEY-----&quot;&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;unwrap&lt;/span&gt;&lt;span&gt;(),&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;Regex&lt;/span&gt;&lt;span&gt;::&lt;/span&gt;&lt;span&gt;new&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;r&quot;AKIA[0-9A-Z]{16}&quot;&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;unwrap&lt;/span&gt;&lt;span&gt;(),&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;Regex&lt;/span&gt;&lt;span&gt;::&lt;/span&gt;&lt;span&gt;new&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;r&quot;(?i)sk-[a-zA-Z0-9]{20,}&quot;&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;unwrap&lt;/span&gt;&lt;span&gt;(),&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;Regex&lt;/span&gt;&lt;span&gt;::&lt;/span&gt;&lt;span&gt;new&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;r&quot;gh[pousr]-[a-zA-Z0-9]{36}&quot;&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;unwrap&lt;/span&gt;&lt;span&gt;(),&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;Regex&lt;/span&gt;&lt;span&gt;::&lt;/span&gt;&lt;span&gt;new&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;r#&quot;(?i)(api_key|secret|token|password)\s*[:=]\s*[&quot;&apos;][a-zA-Z0-9]{32,}[&quot;&apos;]&quot;#&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;unwrap&lt;/span&gt;&lt;span&gt;(),&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;]);&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;div&gt;&lt;div&gt;&lt;/div&gt;&lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;These target the shapes that leak most: RSA and other PEM private key headers, AWS access keys (the &lt;code&gt;AKIA&lt;/code&gt; prefix), OpenAI and Stripe style &lt;code&gt;sk-&lt;/code&gt; keys, GitHub tokens across their prefix family (&lt;code&gt;ghp&lt;/code&gt;, &lt;code&gt;gho&lt;/code&gt;, &lt;code&gt;ghu&lt;/code&gt;, &lt;code&gt;ghs&lt;/code&gt;, &lt;code&gt;ghr&lt;/code&gt;), and the generic &lt;code&gt;api_key = &quot;...&quot;&lt;/code&gt; assignment pattern. Anything matching gets replaced with &lt;code&gt;[REDACTED SECRET]&lt;/code&gt;. This is pattern matching, not proof, so the tool tells you to review output before sharing it. But it means the common ways a key escapes are covered by default, without the user thinking about it.&lt;/p&gt;
&lt;h2&gt;Deciding what is even worth including&lt;/h2&gt;
&lt;p&gt;Before a file’s bytes matter, the tool has to decide whether the file belongs at all. A few cheap filters do most of the work.&lt;/p&gt;
&lt;p&gt;Empty files are dropped on a metadata check before any read. Binary files are caught by sampling: I read the file, look at up to the first 8192 bytes, and if any of them is a null byte, I treat it as binary and skip it.&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;fn&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;is_binary&lt;/span&gt;&lt;span&gt;(content&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;&amp;amp;&lt;/span&gt;&lt;span&gt;[&lt;/span&gt;&lt;span&gt;u8&lt;/span&gt;&lt;span&gt;]) &lt;/span&gt;&lt;span&gt;-&amp;gt;&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;bool&lt;/span&gt;&lt;span&gt; {&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;let&lt;/span&gt;&lt;span&gt; len &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;std&lt;/span&gt;&lt;span&gt;::&lt;/span&gt;&lt;span&gt;cmp&lt;/span&gt;&lt;span&gt;::&lt;/span&gt;&lt;span&gt;min&lt;/span&gt;&lt;span&gt;(content&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;len&lt;/span&gt;&lt;span&gt;(), &lt;/span&gt;&lt;span&gt;8192&lt;/span&gt;&lt;span&gt;);&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;content[&lt;/span&gt;&lt;span&gt;..&lt;/span&gt;&lt;span&gt;len]&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;contains&lt;/span&gt;&lt;span&gt;(&lt;/span&gt;&lt;span&gt;&amp;amp;&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;}&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;div&gt;&lt;div&gt;&lt;/div&gt;&lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;That heuristic is crude and it is exactly right for this job. Real source code effectively never contains a null byte in its first 8KB, and images, compiled objects, and archives almost always do. Whitespace-only files are dropped after read, since a file that trims to nothing adds zero signal and non-zero tokens.&lt;/p&gt;
&lt;p&gt;Large files get a different treatment. Anything over 1MB would dominate the budget, so instead of including or dropping it wholesale, the tool keeps the first 50 and last 50 lines and drops the middle with a marker noting how many lines were omitted. You usually want the imports and the shape of a big generated file, not its ten thousand middle lines.&lt;/p&gt;
&lt;h2&gt;Token accounting so you know before you paste&lt;/h2&gt;
&lt;p&gt;Every artifact carries a token estimate, and the tool sums them into the header. The estimate is deliberately simple: characters divided by four.&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;const&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;CHARS_PER_TOKEN&lt;/span&gt;&lt;span&gt;:&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;usize&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;4&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;let&lt;/span&gt;&lt;span&gt; token_estimate &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; content_str&lt;/span&gt;&lt;span&gt;.&lt;/span&gt;&lt;span&gt;len&lt;/span&gt;&lt;span&gt;() &lt;/span&gt;&lt;span&gt;/&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;CHARS_PER_TOKEN&lt;/span&gt;&lt;span&gt;;&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;div&gt;&lt;div&gt;&lt;/div&gt;&lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;This is not a real tokenizer, and it does not try to be. Pulling in a model-specific BPE tokenizer would add a heavy dependency, tie the output to one model’s vocabulary, and slow the whole thing down for a number that is a budgeting hint, not a billing figure. The four-characters-per-token approximation is close enough to tell you at a glance whether the output fits in a context window, which is the only decision the number needs to support.&lt;/p&gt;
&lt;h2&gt;Parallelism, and streaming instead of buffering&lt;/h2&gt;
&lt;p&gt;File discovery is sequential because the directory walk is inherently ordered, but reading and processing every file is embarrassingly parallel. That phase runs across all cores with rayon: &lt;code&gt;collected_paths.par_iter()&lt;/code&gt; fans the per-file work out, and a progress bar increments as each file lands. On a large repo this is where the wall-clock time is won.&lt;/p&gt;
&lt;p&gt;Output is streamed, not assembled in memory. Rather than building one giant string and writing it at the end, the tool writes each artifact directly through a &lt;code&gt;BufWriter&lt;/code&gt; as it goes. That keeps memory flat even on huge repositories, since the peak footprint is the artifacts themselves plus a small buffer, not a second full copy of the concatenated output.&lt;/p&gt;
&lt;h2&gt;The output format is tuned for how models read code&lt;/h2&gt;
&lt;p&gt;The default output is Markdown, with JSON and XML available. Markdown is the default for a reason: it is what these models were trained on the most, and fenced code blocks with a language hint (&lt;code&gt;rust, &lt;/code&gt;python) are the strongest signal you can give a model about where one file ends and the next begins.&lt;/p&gt;
&lt;p&gt;The document leads with a header line that states the file count and total token estimate, then a &lt;code&gt;text&lt;/code&gt; fenced project tree so the model sees the structure before the contents, then each file as its own section with a metadata line (language, line count, token estimate, and a truncation flag when relevant) above its fenced body. JSON and XML exist for programmatic consumers, and the XML path carefully escapes &lt;code&gt;&amp;amp;&lt;/code&gt;, &lt;code&gt;&amp;lt;&lt;/code&gt;, &lt;code&gt;&amp;gt;&lt;/code&gt;, quotes, and apostrophes so file contents cannot break the document.&lt;/p&gt;
&lt;h2&gt;Why Rust was the right call&lt;/h2&gt;
&lt;p&gt;This tool touches the filesystem hard, needs to be safe by default, and wants to fan work across cores. Rust gives me all three without compromise: the &lt;code&gt;ignore&lt;/code&gt; crate for correct traversal, rayon for parallelism that is a one-line change, and a compiler that will not let me leak a buffer or race a shared regex table. The result starts instantly, holds flat memory on repos of any size, and finishes fast enough that the token count is printed before you have finished reading the command you typed.&lt;/p&gt;
&lt;p&gt;It is MIT licensed and it is the open-source project I reach for most, because the alternative is pasting files into a chat one at a time and hoping I did not include the wrong one.&lt;/p&gt;</content:encoded><category>Rust</category><category>Developer Tools</category><category>LLM</category><category>CLI</category><category>Open Source</category></item><item><title>Fast, Near-Lossless CPU OCR: Running PP-OCRv6 on OpenVINO</title><link>https://irvingernesto.com/blog/fast-lossless-cpu-ocr-with-openvino/</link><guid isPermaLink="true">https://irvingernesto.com/blog/fast-lossless-cpu-ocr-with-openvino/</guid><description>How I made PP-OCRv6 text detection and recognition run 1.4 to 2.7x faster on a plain CPU with OpenVINO, why ONNX Runtime alone barely helps, the exact accuracy cost of the one speed trick that is not free, and the benchmark I built to prove all of it.</description><pubDate>Tue, 20 Jan 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;OCR is one of those workloads that quietly pushes you toward a GPU bill. You want to pull text out of receipts, forms, or scanned pages, you reach for a good model, and the fast path everyone shows you runs on a graphics card. But most people running OCR at volume are running it on servers that already have CPUs sitting mostly idle. The GPU is there to make the model fast, not because the task demands one. So I asked a narrower question: how fast can I make &lt;a href=&quot;https://huggingface.co/PaddlePaddle&quot;&gt;PP-OCRv6&lt;/a&gt;, a genuinely strong open OCR model, on an ordinary CPU, without touching the weights and without lying about what it costs in accuracy?&lt;/p&gt;
&lt;p&gt;The answer is &lt;a href=&quot;https://github.com/Sekinal/ppocrv6-fast-cpu&quot;&gt;ppocrv6-fast-cpu&lt;/a&gt;: the unchanged PP-OCRv6 model, detection plus recognition, run through &lt;a href=&quot;https://github.com/openvinotoolkit/openvino&quot;&gt;OpenVINO&lt;/a&gt;. On my CPU it comes out 1.4x faster than native PyTorch in &lt;code&gt;exact&lt;/code&gt; mode and up to 2.7x in &lt;code&gt;fast&lt;/code&gt; (bf16). The spine of this whole project is a pair of words that usually do not go together: fast AND honest about the tiny accuracy cost. Here is what that actually meant to build.&lt;/p&gt;
&lt;h2&gt;OpenVINO is the win, not ONNX&lt;/h2&gt;
&lt;p&gt;The obvious first move for CPU inference is to export the model to ONNX and run ONNX Runtime. Everyone does this. So I did it, measured it, and it was almost worthless: plain ONNX Runtime on CPU came in at 1.04x on the small tier and 1.11x on the medium tier over PyTorch. That is a rounding error. If I had quoted “we exported to ONNX for CPU speed” as a result, I would have been quoting nothing.&lt;/p&gt;
&lt;p&gt;The real speedup lives one layer down, in the kernels. OpenVINO ships oneDNN AVX-512 convolution kernels that are simply better at saturating a modern CPU than what PyTorch or the ONNX Runtime CPU provider reach for by default. Same model, same math, same inputs: swapping only the engine takes the medium tier from 7385 ms in PyTorch to 5126 ms in OpenVINO &lt;code&gt;exact&lt;/code&gt;, a 1.44x win, and with bf16 down to 2779 ms, 2.66x. The pipeline itself is thin. Detection runs a DB (differentiable binarization) model, I crop the detected text lines, and recognition runs a CTC head over each crop. All the heavy convolution goes to OpenVINO; everything else is deliberately untouched.&lt;/p&gt;
&lt;p&gt;That “deliberately untouched” part is the whole trick to staying accurate. I reuse the exact HuggingFace PP-OCRv6 image processors for pre and post processing, verbatim:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;def&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;processor&lt;/span&gt;&lt;span&gt;(model_id):&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;if&lt;/span&gt;&lt;span&gt; model_id &lt;/span&gt;&lt;span&gt;not&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;in&lt;/span&gt;&lt;span&gt; _proc_cache:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;        &lt;/span&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; transformers &lt;/span&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; AutoImageProcessor&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;        &lt;/span&gt;&lt;/span&gt;&lt;span&gt;_proc_cache[model_id] &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; AutoImageProcessor.from_pretrained(model_id)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;return&lt;/span&gt;&lt;span&gt; _proc_cache[model_id]&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;div&gt;&lt;div&gt;&lt;/div&gt;&lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;The only thing that changes versus native PP-OCRv6 is how the convolutions execute. Feed OpenVINO fp32 the identical input, and its output is bit-identical to PyTorch. I confirmed that on 34 of 34 documents. At the engine level, the speed is free.&lt;/p&gt;
&lt;h2&gt;The one optimization that is not free, and I say so&lt;/h2&gt;
&lt;p&gt;If the engine is bit-identical, why do I call the whole thing near-lossless instead of lossless? Because of one pipeline decision, and it is the interesting one.&lt;/p&gt;
&lt;p&gt;Recognition runs once per detected text line, and text crops vary wildly in width: a two-character label and a ninety-character sentence are the same model, different input shapes. The naive fix is to pad every crop in a batch to a common width. That is a trap: padding to a shared width corrupts the CTC decode, because the recognizer reads the padded region as content and the transcription drifts. So I process each crop at its natural width instead. Correct, but now OpenVINO sees a new input shape on nearly every crop and recompiles, which thrashes.&lt;/p&gt;
&lt;p&gt;The compromise is width-bucketing. I round each crop’s width up to the next multiple of 64, so OpenVINO only ever sees a handful of shapes and caches them:&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;w &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; pv.shape[&lt;/span&gt;&lt;span&gt;3&lt;/span&gt;&lt;span&gt;]&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;wb &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; ((w &lt;/span&gt;&lt;span&gt;+&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;_REC_WMULT&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;-&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;1&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;//&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;_REC_WMULT&lt;/span&gt;&lt;span&gt;) &lt;/span&gt;&lt;span&gt;*&lt;/span&gt;&lt;span&gt; &lt;/span&gt;&lt;span&gt;_REC_WMULT&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;if&lt;/span&gt;&lt;span&gt; wb &lt;/span&gt;&lt;span&gt;&amp;gt;&lt;/span&gt;&lt;span&gt; w:&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;&lt;span&gt;    &lt;/span&gt;&lt;/span&gt;&lt;span&gt;pv &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; np.pad(pv, ((&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;), (&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;), (&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;, &lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;), (&lt;/span&gt;&lt;span&gt;0&lt;/span&gt;&lt;span&gt;, wb &lt;/span&gt;&lt;span&gt;-&lt;/span&gt;&lt;span&gt; w)))&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;div&gt;&lt;div&gt;&lt;/div&gt;&lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;This is worth roughly a 1.4x on recognition and it fixes the shape thrash. But it is the one place the output can differ from native, because native runs every crop at its exact natural width and I run it at the bucketed width. That residual is small: &lt;code&gt;exact&lt;/code&gt; mode is over 99.6% character-identical to native, 0.09% CER on the small tier and 0.36% on the medium tier. It is not zero. I could make it zero by disabling bucketing and eating the speed loss, and the code lets you do exactly that. But I refuse to call this bit-lossless when it is not, because the difference between “lossless” and “99.9% lossless with a named, understood cause” is the difference between marketing and engineering.&lt;/p&gt;
&lt;p&gt;There are two other honest wins worth naming. The &lt;code&gt;fast&lt;/code&gt; mode runs bf16, which roughly doubles throughput on AVX-512-BF16 CPUs; detection boxes stay identical because the DB post-process is robust to the rounding, and recognition takes about 0.5% CER, with the rare misses landing on the hardest glyphs like stylized fonts and some CJK punctuation. And on the medium detector I do a lossless structural reparameterization: RepLKFPN’s IntraclassBlocks sum three parallel convolutions per group (a symmetric KxK plus a vertical and a horizontal strip), and because they share input and output shape they fuse exactly into one KxK conv by zero-padding the smaller kernels and summing weights. Nine convs per block become three, and the output is identical up to float accumulation order. That one is genuinely free.&lt;/p&gt;
&lt;h2&gt;Why the benchmark is the actual product&lt;/h2&gt;
&lt;p&gt;Here is the part I care about most. Anyone can run a model twice and tweet a speedup. I did not want a number, I wanted a result I would trust if someone else published it, so the benchmark is built to survive scrutiny.&lt;/p&gt;
&lt;p&gt;I measured 143 documents: 131 dense arXiv pages rendered at 150 DPI, deliberately text-heavy at 40 to 250 lines each, plus 12 scene and multilingual images to stress detection on non-document layouts. Every configuration, two tiers by two modes, ran against native PyTorch, plain ONNX Runtime, and off-the-shelf RapidOCR. Point estimates carry 95% confidence intervals: t-based for means, bootstrap with 5000 resamples for medians and CER, Clopper-Pearson exact intervals for the page-match proportion. Engine and mode contrasts are tested with paired Wilcoxon signed-rank tests, because latency is right-skewed and a non-parametric paired test is the honest choice. The figures are R and ggplot2.&lt;/p&gt;
&lt;p&gt;That rigor is not decoration, it is what lets me make the honesty precise instead of hand-wavy. It is why I can tell you the sub-0.4% CER on &lt;code&gt;exact&lt;/code&gt; comes from bucketing and not from OpenVINO, because the engine-only comparison shows bit-identical output while the full pipeline shows the residual. It is why I can say latency is driven by text density and not page count: the distribution is bimodal, sparse scene images sit near 50 to 450 ms while dense pages sit at 1 to 6 seconds, and the same medium &lt;code&gt;exact&lt;/code&gt; config that takes 6 seconds on a 50-line arXiv page takes 457 ms median on the scene images. Plan capacity by characters per page, not pages per second. It is also why I can report where I lose: RapidOCR runs older PP-OCRv4 models, so that is a tool comparison and not a same-model engine swap, and I label it as such rather than pretending it is apples to apples.&lt;/p&gt;
&lt;p&gt;The measurement rules are boring on purpose. Latency is the minimum of two timed runs per document, so the warm second run rejects shape-compilation spikes and scheduler noise. Accuracy is measured against each tier’s own native PyTorch output computed with identical pre and post processing, a faithful-reproduction metric, not human ground truth, so the only variable is the inference path. Hardware is an AMD Ryzen 7 8845HS, Zen 4, 8 physical cores with AVX-512-BF16, running 8 threads, and I say that loudly because absolute latency is hardware-dependent even though the relative engine, mode, and tier results are stable.&lt;/p&gt;
&lt;h2&gt;What this is, and what it is not&lt;/h2&gt;
&lt;p&gt;This wraps the public, Apache-2.0 PP-OCRv6 model unchanged. There is no custom-trained model here and no new weights. All the speed comes from the inference path: the right engine, a lossless reparameterization, bf16 where it is safe, and a width-bucketing knob whose exact cost I can quote to two decimal places. Self-hosting it wins on cost at volume, on privacy because no data leaves your network, and on running air-gapped, at the price of running it yourself.&lt;/p&gt;
&lt;p&gt;I think the honesty is the feature. It is easy to ship “lossless CPU OCR, 2.7x faster.” It is harder, and more useful, to ship “1.4 to 2.7x faster, over 99.6% character-identical, here is the one place it is not perfect and exactly why, and here are 143 documents with confidence intervals so you can check me.” The second one is the one I would want to depend on.&lt;/p&gt;</content:encoded><category>OCR</category><category>OpenVINO</category><category>Performance</category><category>CPU</category><category>Open Source</category></item><item><title>NeuralTranslate: Preserving Nahuatl with AI</title><link>https://irvingernesto.com/blog/neuraltranslate-preserving-nahuatl-with-ai/</link><guid isPermaLink="true">https://irvingernesto.com/blog/neuraltranslate-preserving-nahuatl-with-ai/</guid><description>I fine-tuned Gemma 3 27B for Nahuatl to Spanish and hit 97.48 ChrF on validation. The number was real, and still misleading, and Nahuatl speakers are the ones who showed me why.</description><pubDate>Fri, 20 Jun 2025 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Nahuatl, the language of the Aztecs, is still spoken by over 1.7 million people in Mexico today. Yet it remains severely underrepresented in modern NLP systems. When I started this project, most translation tools either ignored Nahuatl entirely or produced unusable results.&lt;/p&gt;
&lt;p&gt;I set out to change that.&lt;/p&gt;
&lt;h2&gt;The challenge&lt;/h2&gt;
&lt;p&gt;Building a neural machine translation system for Nahuatl presents unique challenges:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Low-resource language&lt;/strong&gt;: unlike Spanish or English, there’s limited parallel corpus data available&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dialectal variation&lt;/strong&gt;: Nahuatl has many regional variants (Classical, Huasteca, Central…)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Complex morphology&lt;/strong&gt;: Nahuatl is polysynthetic: single words can express what takes entire sentences in English&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Computational constraints&lt;/strong&gt;: state-of-the-art models require massive GPU resources&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I needed a model powerful enough to understand Nahuatl’s complexity, yet efficient enough to train on available hardware.&lt;/p&gt;
&lt;h2&gt;Why Gemma 3 27B?&lt;/h2&gt;
&lt;p&gt;After benchmarking across the Gemma 3 family, the 27B parameter model stood out. Smaller variants (1B, 4B) lacked the capacity to capture Nahuatl’s morphological richness, while 27B showed promising zero-shot understanding of linguistic patterns.&lt;/p&gt;
&lt;p&gt;The problem? A 27B model in full precision needs ~108GB of VRAM just for weights, before optimizer states, gradients, and activations.&lt;/p&gt;
&lt;h2&gt;The solution: 4-bit quantization + full fine-tune&lt;/h2&gt;
&lt;p&gt;Here’s where it gets interesting. Instead of using QLoRA (which freezes most weights and only trains adapter layers), I went for a &lt;strong&gt;full fine-tune on 4-bit quantized weights&lt;/strong&gt;.&lt;/p&gt;
&lt;div&gt;&lt;figure&gt;&lt;figcaption&gt;&lt;/figcaption&gt;&lt;pre&gt;&lt;code&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;from&lt;/span&gt;&lt;span&gt; transformers &lt;/span&gt;&lt;span&gt;import&lt;/span&gt;&lt;span&gt; BitsAndBytesConfig&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;
&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;bnb_config &lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt; BitsAndBytesConfig(&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;load_in_4bit&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;True&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;bnb_4bit_quant_type&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;&quot;nf4&quot;&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;bnb_4bit_compute_dtype&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;torch.bfloat16,&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;    &lt;/span&gt;&lt;span&gt;bnb_4bit_use_double_quant&lt;/span&gt;&lt;span&gt;=&lt;/span&gt;&lt;span&gt;True&lt;/span&gt;&lt;span&gt;,&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;div&gt;&lt;div&gt;&lt;span&gt;)&lt;/span&gt;&lt;/div&gt;&lt;/div&gt;&lt;/code&gt;&lt;/pre&gt;&lt;div&gt;&lt;div&gt;&lt;/div&gt;&lt;div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/figure&gt;&lt;/div&gt;
&lt;p&gt;This approach:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Reduces the memory footprint&lt;/strong&gt; from ~108GB to ~27GB for weights&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Maintains full model plasticity&lt;/strong&gt;, unlike LoRA/QLoRA&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fits on a single B200&lt;/strong&gt; (180GB VRAM)&lt;/li&gt;
&lt;/ul&gt;
&lt;aside&gt; &lt;p&gt;Note&lt;/p&gt; &lt;div&gt; &lt;p&gt;The key insight: QLoRA is great for adapting models to new tasks while preserving general
knowledge. But for low-resource language translation, the model needs to fundamentally
restructure its internal representations, which requires updating all parameters.&lt;/p&gt; &lt;/div&gt; &lt;/aside&gt; 
&lt;h2&gt;Training setup&lt;/h2&gt;





































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Parameter&lt;/th&gt;&lt;th&gt;Value&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Base model&lt;/td&gt;&lt;td&gt;&lt;code&gt;google/gemma-3-27b-it&lt;/code&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Quantization&lt;/td&gt;&lt;td&gt;4-bit NF4&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Learning rate&lt;/td&gt;&lt;td&gt;2e-5&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Batch size&lt;/td&gt;&lt;td&gt;4 (gradient accumulation)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Epochs&lt;/td&gt;&lt;td&gt;~3&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hardware&lt;/td&gt;&lt;td&gt;NVIDIA B200 (180GB VRAM)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Training time&lt;/td&gt;&lt;td&gt;~8 hours&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;I trained on the SomosNLP Axolotl Nahuatl-Spanish parallel corpus, an existing open dataset. That choice turned out to matter more than any hyperparameter, for a reason I get to below.&lt;/p&gt;
&lt;h2&gt;Results&lt;/h2&gt;
&lt;p&gt;The training curves tell the story better than any table:&lt;/p&gt;
&lt;figure&gt; &lt;figcaption&gt;Training run: eval loss vs ChrF, Gemma 3 27B&lt;/figcaption&gt;    60   70   80   90   100  5001500250035004355 ChrF loss (log)    97.48 ChrF  &lt;div&gt; &lt;span&gt;&lt;span&gt;&lt;/span&gt;ChrF score&lt;/span&gt; &lt;span&gt;&lt;span&gt;&lt;/span&gt;Eval loss&lt;/span&gt; &lt;/div&gt; &lt;/figure&gt;
&lt;p&gt;By step 4,355 the model reached a &lt;strong&gt;ChrF of 97.48 on the validation set&lt;/strong&gt;. On paper that looks like a solved problem. It is not, and the distance between that number and the truth is the most important thing this project taught me. I get to it two sections down.&lt;/p&gt;
&lt;aside&gt; &lt;p&gt;Note&lt;/p&gt; &lt;div&gt; &lt;p&gt;For a technical comparison on the same validation data: standard QLoRA plateaued around 85 to 88
ChrF, and the full fine-tune gained nearly 10 points. That gap is real. It is a statement about the
training method, not about how good a modern Nahuatl translator the result actually is.&lt;/p&gt; &lt;/div&gt; &lt;/aside&gt; 
&lt;h2&gt;Why ChrF?&lt;/h2&gt;
&lt;p&gt;I chose ChrF (character n-gram F-score) over BLEU because:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Morphologically rich languages&lt;/strong&gt;: ChrF handles agglutinative languages better by comparing character sequences rather than word tokens&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No tokenization dependency&lt;/strong&gt;: Nahuatl lacks standardized tokenization, making word-based metrics unreliable&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Better correlation with human judgment&lt;/strong&gt; for morphologically complex languages&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;Example translations&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Nahuatl&lt;/strong&gt;: &lt;em&gt;Nimitztlazohtla&lt;/em&gt; → &lt;strong&gt;Model&lt;/strong&gt;: &lt;em&gt;Te amo&lt;/em&gt; (I love you)&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nahuatl&lt;/strong&gt;: &lt;em&gt;Tlein motocah?&lt;/em&gt; → &lt;strong&gt;Model&lt;/strong&gt;: &lt;em&gt;¿Cómo te llamas?&lt;/em&gt; (What is your name?)&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Nahuatl&lt;/strong&gt;: &lt;em&gt;Nicnequi atl&lt;/em&gt; → &lt;strong&gt;Model&lt;/strong&gt;: &lt;em&gt;Quiero agua&lt;/em&gt; (I want water)&lt;/p&gt;
&lt;p&gt;The model correctly handles incorporated pronouns (&lt;em&gt;ni-&lt;/em&gt; = I, &lt;em&gt;mitz-&lt;/em&gt; = you), verb conjugations and aspects, and interrogative particles.&lt;/p&gt;
&lt;h2&gt;The validation score was real, and still misleading&lt;/h2&gt;
&lt;p&gt;A ChrF of 97 on validation does not mean the model is a good modern Nahuatl translator. It means the model learned to reproduce the validation text very faithfully. Everything then depends on what that text was.&lt;/p&gt;
&lt;p&gt;The Axolotl corpus uses an older orthography that does not match how Nahuatl is written by modern translators today. So the model became excellent at matching flawed-orthography references, and ChrF, which only compares character sequences against those references, happily rewarded it. A metric measures agreement with your data. If the data is systematically off, a high score just means you learned the flaw well.&lt;/p&gt;
&lt;p&gt;The people who caught this were not a benchmark. Nahuatl speakers read the actual output and told me it was wrong for modern Nahuatl. That feedback is worth more than the 97, and it is what pushed me to go back and take the orthography and the dataset seriously rather than trusting the curve. The lesson I keep: for a language you do not speak, the real evaluation is the community, not the validation set.&lt;/p&gt;
&lt;h2&gt;Try it yourself&lt;/h2&gt;
&lt;a href=&quot;https://huggingface.co/spaces/Thermostatic/neuraltranslate-27b-mt-nah-es&quot; target=&quot;_blank&quot;&gt; &lt;span&gt; &lt;span&gt;huggingface.co&lt;/span&gt; &lt;span&gt; NeuralTranslate Live Demo &lt;/span&gt; &lt;span&gt;Try the Nahuatl-Spanish translator in your browser&lt;/span&gt; &lt;/span&gt; &lt;span&gt;    &lt;/span&gt; &lt;/a&gt;
&lt;a href=&quot;https://huggingface.co/Thermostatic&quot; target=&quot;_blank&quot;&gt; &lt;span&gt; &lt;span&gt;huggingface.co&lt;/span&gt; &lt;span&gt; Models on HuggingFace &lt;/span&gt; &lt;span&gt;Download the fine-tuned weights, plus the rest of the low-resource language line&lt;/span&gt; &lt;/span&gt; &lt;span&gt;    &lt;/span&gt; &lt;/a&gt;
&lt;h2&gt;What’s next&lt;/h2&gt;
&lt;p&gt;This project taught me that the model was the easy part and the data and the community were the hard, important parts. Since this post was written, the effort has grown into the &lt;strong&gt;Rosettia&lt;/strong&gt; family, including a &lt;a href=&quot;/blog/rosettia-low-resource-languages&quot;&gt;Spanish to Quechua system&lt;/a&gt; that beat the prior task winners, alongside speech recognition for Mexican Indigenous languages.&lt;/p&gt;
&lt;hr /&gt;
&lt;p&gt;&lt;em&gt;This project is dedicated to preserving indigenous languages through technology. If you’re working on similar efforts, I’d love to connect.&lt;/em&gt;&lt;/p&gt;</content:encoded><category>AI</category><category>NLP</category><category>Fine-tuning</category><category>Nahuatl</category><category>LLM</category></item></channel></rss>