Skip to content

← Writing / June 20, 2025 / 5 min read

NeuralTranslate: Preserving Nahuatl with AI

I fine-tuned Gemma 3 27B for Nahuatl to Spanish and hit 97.48 ChrF on validation. The number was real, and still misleading, and Nahuatl speakers are the ones who showed me why.

AINLPFine-tuningNahuatlLLM

Nahuatl, the language of the Aztecs, is still spoken by over 1.7 million people in Mexico today. Yet it remains severely underrepresented in modern NLP systems. When I started this project, most translation tools either ignored Nahuatl entirely or produced unusable results.

I set out to change that.

The challenge

Building a neural machine translation system for Nahuatl presents unique challenges:

  1. Low-resource language: unlike Spanish or English, there’s limited parallel corpus data available
  2. Dialectal variation: Nahuatl has many regional variants (Classical, Huasteca, Central…)
  3. Complex morphology: Nahuatl is polysynthetic: single words can express what takes entire sentences in English
  4. Computational constraints: state-of-the-art models require massive GPU resources

I needed a model powerful enough to understand Nahuatl’s complexity, yet efficient enough to train on available hardware.

Why Gemma 3 27B?

After benchmarking across the Gemma 3 family, the 27B parameter model stood out. Smaller variants (1B, 4B) lacked the capacity to capture Nahuatl’s morphological richness, while 27B showed promising zero-shot understanding of linguistic patterns.

The problem? A 27B model in full precision needs ~108GB of VRAM just for weights, before optimizer states, gradients, and activations.

The solution: 4-bit quantization + full fine-tune

Here’s where it gets interesting. Instead of using QLoRA (which freezes most weights and only trains adapter layers), I went for a full fine-tune on 4-bit quantized weights.

from transformers import BitsAndBytesConfig
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True,
)

This approach:

  • Reduces the memory footprint from ~108GB to ~27GB for weights
  • Maintains full model plasticity, unlike LoRA/QLoRA
  • Fits on a single B200 (180GB VRAM)

Training setup

ParameterValue
Base modelgoogle/gemma-3-27b-it
Quantization4-bit NF4
Learning rate2e-5
Batch size4 (gradient accumulation)
Epochs~3
HardwareNVIDIA B200 (180GB VRAM)
Training time~8 hours

I trained on the SomosNLP Axolotl Nahuatl-Spanish parallel corpus, an existing open dataset. That choice turned out to matter more than any hyperparameter, for a reason I get to below.

Results

The training curves tell the story better than any table:

Training run: eval loss vs ChrF, Gemma 3 27B
60 70 80 90 100 5001500250035004355 ChrF loss (log) 97.48 ChrF
ChrF score Eval loss

By step 4,355 the model reached a ChrF of 97.48 on the validation set. On paper that looks like a solved problem. It is not, and the distance between that number and the truth is the most important thing this project taught me. I get to it two sections down.

Why ChrF?

I chose ChrF (character n-gram F-score) over BLEU because:

  1. Morphologically rich languages: ChrF handles agglutinative languages better by comparing character sequences rather than word tokens
  2. No tokenization dependency: Nahuatl lacks standardized tokenization, making word-based metrics unreliable
  3. Better correlation with human judgment for morphologically complex languages

Example translations

Nahuatl: Nimitztlazohtla → Model: Te amo (I love you)

Nahuatl: Tlein motocah? → Model: ¿Cómo te llamas? (What is your name?)

Nahuatl: Nicnequi atl → Model: Quiero agua (I want water)

The model correctly handles incorporated pronouns (ni- = I, mitz- = you), verb conjugations and aspects, and interrogative particles.

The validation score was real, and still misleading

A ChrF of 97 on validation does not mean the model is a good modern Nahuatl translator. It means the model learned to reproduce the validation text very faithfully. Everything then depends on what that text was.

The Axolotl corpus uses an older orthography that does not match how Nahuatl is written by modern translators today. So the model became excellent at matching flawed-orthography references, and ChrF, which only compares character sequences against those references, happily rewarded it. A metric measures agreement with your data. If the data is systematically off, a high score just means you learned the flaw well.

The people who caught this were not a benchmark. Nahuatl speakers read the actual output and told me it was wrong for modern Nahuatl. That feedback is worth more than the 97, and it is what pushed me to go back and take the orthography and the dataset seriously rather than trusting the curve. The lesson I keep: for a language you do not speak, the real evaluation is the community, not the validation set.

Try it yourself

huggingface.co NeuralTranslate Live Demo Try the Nahuatl-Spanish translator in your browser huggingface.co Models on HuggingFace Download the fine-tuned weights, plus the rest of the low-resource language line

What’s next

This project taught me that the model was the easy part and the data and the community were the hard, important parts. Since this post was written, the effort has grown into the Rosettia family, including a Spanish to Quechua system that beat the prior task winners, alongside speech recognition for Mexican Indigenous languages.


This project is dedicated to preserving indigenous languages through technology. If you’re working on similar efforts, I’d love to connect.