← Writing / June 20, 2025 / 5 min read
NeuralTranslate: Preserving Nahuatl with AI
I fine-tuned Gemma 3 27B for Nahuatl to Spanish and hit 97.48 ChrF on validation. The number was real, and still misleading, and Nahuatl speakers are the ones who showed me why.
Nahuatl, the language of the Aztecs, is still spoken by over 1.7 million people in Mexico today. Yet it remains severely underrepresented in modern NLP systems. When I started this project, most translation tools either ignored Nahuatl entirely or produced unusable results.
I set out to change that.
The challenge
Building a neural machine translation system for Nahuatl presents unique challenges:
- Low-resource language: unlike Spanish or English, there’s limited parallel corpus data available
- Dialectal variation: Nahuatl has many regional variants (Classical, Huasteca, Central…)
- Complex morphology: Nahuatl is polysynthetic: single words can express what takes entire sentences in English
- Computational constraints: state-of-the-art models require massive GPU resources
I needed a model powerful enough to understand Nahuatl’s complexity, yet efficient enough to train on available hardware.
Why Gemma 3 27B?
After benchmarking across the Gemma 3 family, the 27B parameter model stood out. Smaller variants (1B, 4B) lacked the capacity to capture Nahuatl’s morphological richness, while 27B showed promising zero-shot understanding of linguistic patterns.
The problem? A 27B model in full precision needs ~108GB of VRAM just for weights, before optimizer states, gradients, and activations.
The solution: 4-bit quantization + full fine-tune
Here’s where it gets interesting. Instead of using QLoRA (which freezes most weights and only trains adapter layers), I went for a full fine-tune on 4-bit quantized weights.
from transformers import BitsAndBytesConfig
bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True,)This approach:
- Reduces the memory footprint from ~108GB to ~27GB for weights
- Maintains full model plasticity, unlike LoRA/QLoRA
- Fits on a single B200 (180GB VRAM)
Training setup
| Parameter | Value |
|---|---|
| Base model | google/gemma-3-27b-it |
| Quantization | 4-bit NF4 |
| Learning rate | 2e-5 |
| Batch size | 4 (gradient accumulation) |
| Epochs | ~3 |
| Hardware | NVIDIA B200 (180GB VRAM) |
| Training time | ~8 hours |
I trained on the SomosNLP Axolotl Nahuatl-Spanish parallel corpus, an existing open dataset. That choice turned out to matter more than any hyperparameter, for a reason I get to below.
Results
The training curves tell the story better than any table:
By step 4,355 the model reached a ChrF of 97.48 on the validation set. On paper that looks like a solved problem. It is not, and the distance between that number and the truth is the most important thing this project taught me. I get to it two sections down.
Why ChrF?
I chose ChrF (character n-gram F-score) over BLEU because:
- Morphologically rich languages: ChrF handles agglutinative languages better by comparing character sequences rather than word tokens
- No tokenization dependency: Nahuatl lacks standardized tokenization, making word-based metrics unreliable
- Better correlation with human judgment for morphologically complex languages
Example translations
Nahuatl: Nimitztlazohtla → Model: Te amo (I love you)
Nahuatl: Tlein motocah? → Model: ¿Cómo te llamas? (What is your name?)
Nahuatl: Nicnequi atl → Model: Quiero agua (I want water)
The model correctly handles incorporated pronouns (ni- = I, mitz- = you), verb conjugations and aspects, and interrogative particles.
The validation score was real, and still misleading
A ChrF of 97 on validation does not mean the model is a good modern Nahuatl translator. It means the model learned to reproduce the validation text very faithfully. Everything then depends on what that text was.
The Axolotl corpus uses an older orthography that does not match how Nahuatl is written by modern translators today. So the model became excellent at matching flawed-orthography references, and ChrF, which only compares character sequences against those references, happily rewarded it. A metric measures agreement with your data. If the data is systematically off, a high score just means you learned the flaw well.
The people who caught this were not a benchmark. Nahuatl speakers read the actual output and told me it was wrong for modern Nahuatl. That feedback is worth more than the 97, and it is what pushed me to go back and take the orthography and the dataset seriously rather than trusting the curve. The lesson I keep: for a language you do not speak, the real evaluation is the community, not the validation set.
Try it yourself
huggingface.co NeuralTranslate Live Demo Try the Nahuatl-Spanish translator in your browser huggingface.co Models on HuggingFace Download the fine-tuned weights, plus the rest of the low-resource language lineWhat’s next
This project taught me that the model was the easy part and the data and the community were the hard, important parts. Since this post was written, the effort has grown into the Rosettia family, including a Spanish to Quechua system that beat the prior task winners, alongside speech recognition for Mexican Indigenous languages.
This project is dedicated to preserving indigenous languages through technology. If you’re working on similar efforts, I’d love to connect.
Nahuatl is the language the Aztecs spoke, and it is not a museum piece. More than 1.7 million people in Mexico still speak it today. That is a whole city’s worth of people, larger than the population of many countries. And yet, when you reach for a translation app or an AI assistant, Nahuatl is almost always missing. Most tools either ignore it completely or spit out nonsense. This project set out to fix that, by building an AI that can actually translate between Nahuatl and Spanish.
Why does nobody bother?
The short answer is money and data. Modern translation AIs learn by reading enormous piles of example sentences that have already been translated by humans, the same sentence in two languages, lined up side by side. For English and Spanish, the internet is drowning in such examples. For Nahuatl, there is barely a trickle. When the training material is scarce, a language is called low-resource, and big tech companies mostly do not find it worth their while to build tools for it. So if anyone is going to do it, it has to be someone who cares on purpose.
Why Nahuatl is genuinely hard to translate
Beyond the shortage of examples, Nahuatl has a feature that makes it fascinating and difficult. In English or Spanish, you build a thought out of several separate words. In Nahuatl, a single word can carry what would take a whole sentence in English. The word gets built up by gluing pieces together: a piece for “I,” a piece for “you,” a piece for the verb, a piece for the tense, all fused into one long word. Linguists call this polysynthetic. It means the AI cannot just swap word for word. It has to understand how the pieces snap together and come apart.
On top of that, Nahuatl is not one single language but a family of regional varieties: the classical form, the Huasteca variety, the Central variety, and more. They differ enough that mixing them together confuses the AI. And there is a subtler trap hiding in the training data, one I did not see until later, which I come back to at the end.
The trick: shrink the model without lobotomizing it
The project started from a large, powerful, freely available AI model made by Google, called Gemma. It comes in several sizes. The smaller versions turned out to be too simple to grasp Nahuatl’s tangled word-building. The big version, with 27 billion internal settings, had the horsepower. The catch is that the big version is a monster to run: just holding it in memory normally needs about 108 gigabytes of specialized computer memory, which is far more than a single machine usually has.
The clever move was compression. There is a technique that squeezes the model down so it stores its numbers more coarsely, roughly quartering the memory it needs, from about 108 down to about 27 gigabytes, so the whole thing fits on a single (very fancy) chip. The usual way people use this trick is to freeze most of the model and only teach a thin new layer on top. But this project did something bolder: it kept the whole model flexible and retrained all of it. The reasoning is that learning a new language is not a light touch-up. The AI has to genuinely rewire how it thinks, and that means letting every part of it change.
Did it work? Yes, and then no, and that is the real story
On paper, spectacularly. To measure quality, the project used a scoring method that compares translations letter by letter, which is fairer for a language that builds giant words out of small pieces. On a scale where a perfect match is 100, the model scored about 97.5. Training took only about eight hours on one high-end chip. Give it the single Nahuatl word Nimitztlazohtla and it returns “I love you,” give it Tlein motocah? and it returns “What is your name?”
But here is the honest catch, and it is the most important part of the whole project. That 97.5 was measured against the same collection of examples the model learned from, and that collection was written in an older spelling of Nahuatl that does not match how the language is actually written today. So the model became very good at copying an outdated spelling, and the score cheerfully rewarded it for exactly that. A high score against flawed examples does not mean good translation. It means you learned the flaw well.
The people who caught this were not a test. They were Nahuatl speakers who read the real translations and told me, plainly, that they were off for modern Nahuatl. That feedback was worth more than the 97.5, and it is what made me go back and take the spelling and the data seriously instead of trusting the number. The lesson I keep is simple: for a language you do not speak, the real judge is the people who do, not a score.
Why this matters
The translator is free for anyone to try in a browser and free for others to build on, but the bigger takeaway is that lesson. A single determined person can build real language tools for a language the tech giants overlook, as long as they remember that the model is the easy part and the data and the community are the hard, important ones. This project became the seed of a larger effort, later named Rosettia, that went on to other forgotten languages, including a Spanish to Quechua translator that beat the previous record-holders. Languages do not vanish only because people stop speaking them. They vanish when nobody builds the tools that let them live in the modern world. Those tools can be built, carefully, and with the speakers in the room.