← Writing / July 9, 2026 / 6 min read
GSPO: Reinforcement Learning for Low-Resource Speech Recognition
Teaching Whisper to hear Mexican Indigenous languages with RL. Hand-rolling GSPO around an audio policy, rewarding on negative Character Error Rate, and making real progress on languages every off-the-shelf system fails.
Every state-of-the-art speech recognizer can hear Spanish. None of them can hear the languages Spanish landed on top of. Point Whisper, NVIDIA Parakeet or Canary, Meta MMS, or IBM Granite at Mexican Spanish and you get roughly 14% word error rate: usable. Point the exact same models at Nahuatl, or any of the 22 other Indigenous languages of Mexico, and word error rate jumps to 99% or worse. The model is not transcribing. It is guessing in the wrong language.
That gap, 14% versus 99%, is the whole project. It is not a tuning problem. It is a “the training data never existed” problem, spanning 23 languages across six or more language families. This post is about closing part of that gap with reinforcement learning, and about the several ways I nearly fooled myself while doing it.
First you need a benchmark you cannot cheat
Before touching a model I built the evaluation, because in low-resource work the benchmark is where projects quietly die. If you cannot tell progress from noise, every training run looks like progress.
The result is MEXA, a contamination-resistant ASR benchmark. Two pieces make it hard to game:
- A private held-out test set that never ships.
- A public audio-plus-text fingerprint registry, so anyone can decontaminate their own training data against the test without ever seeing the test itself. You check whether your clips collide with the fingerprints; you learn nothing about the answers.
The scoring is linguistically aware. A naive normalizer would strip the tone marks and glottal stops that carry meaning in these languages, quietly rewarding a model for mangling them. MEXA’s normalization preserves those marks through both WER and CER, so the metric measures the language people actually speak, not a flattened ASCII shadow of it.
Supervised first, then reinforcement
The base model is openai/whisper-large-v3-turbo. The pipeline is two stages.
Stage one is a supervised fine-tune: LoRA through Unsloth, with the hyperparameters found by Optuna rather than by me guessing. The winning configuration used rsLoRA plus DoRA at rank 128. This gets the model into the neighborhood of the languages. It stops guessing in Spanish and starts producing plausible orthography.
Stage two is where it gets interesting: reinforcement learning with GSPO, Group Sequence Policy Optimization (arXiv 2507.18071).
Why I had to hand-roll the trainer
The obvious move is to reach for TRL’s GRPOTrainer and be done in an afternoon. It does not work, and the reason is structural: every mainstream RL-for-LLM library assumes text in, text out. The policy is a language model, the prompt is tokens, the completion is tokens.
Whisper’s input is not tokens. It is audio. The policy conditions on a log-mel spectrogram, samples a transcript, and gets scored against a reference. So the entire trainer had to be rebuilt around an audio policy: batching audio, generating groups of candidate transcripts per clip, computing rewards, and pushing the GSPO update, with the encoder in the loop the whole way.
for batch of audio clips: for each clip: sample G candidate transcripts from the policy # audio -> text reward[g] = -CER(candidate[g], reference) # verifiable, no reward model advantages = group_normalize(rewards) # per-clip baseline ratio = exp((logp_theta - logp_old) / len(y)) # sequence-level, length-normalized loss = -(ratio * advantages) + beta * KL(policy || sft_anchor) step(loss) # lr = 1e-6The reward is just negative CER, and that is the point
The reward function is negative Character Error Rate against the reference transcript. That is it. No learned reward model, no human preference data. This is RLVR, reinforcement learning from a verifiable reward: the reference is ground truth, so the reward is exact and free of the reward-hacking pathologies that plague learned critics.
It also has an elegant side effect. Whisper’s worst failure mode on hard audio is hallucination: repetition loops where it emits the same phrase over and over. Under a CER reward those loops punish themselves, because every hallucinated character counts as an insertion error and inflates CER. The model is not told “stop looping.” It discovers that looping is expensive.
To keep the policy from drifting off a cliff, there is a KL penalty (the k3 estimator) back to a frozen SFT anchor, and the learning rate is deliberately tiny at 1e-6. RL here is a scalpel, not a hammer.
GSPO versus GRPO, briefly
The difference from GRPO is the importance ratio. GRPO computes it per token. GSPO computes it once per sequence, length-normalized:
ratio = exp( (log p_theta(y|x) - log p_old(y|x)) / |y| )For long sequences the token-level product in GRPO accumulates variance and can blow up. The sequence-level, length-normalized ratio is far more stable, which matters when your outputs are full transcripts rather than short answers.
Results
On the 240-clip MEXA subset, evaluated fairly through faster-whisper:
| Stage | WER | CER |
|---|---|---|
| SFT preview | 68.5 | 30.6 |
| SFT + GSPO | 66.0 | 28.3 |
A net improvement of 2.45 WER from RL alone. The comparison that actually matters is not my own benchmark, it is the gap to what already exists: off-the-shelf systems (Whisper, Parakeet, Canary, MMS, Granite) sit near or above 99% WER on these languages, effectively unusable, and where the next-best measured system lands at 74.3 CER, this one reaches 30.6. The point is not a scoreboard rank on a benchmark I built, it is that the languages were near-untranscribable and now they are not.
The clearest evidence that the CER reward did what it was designed to do: the worst looping language dropped from 100.2 WER to 75.9. That is the hallucination penalty working exactly as predicted, on the language that needed it most.
The negative result I am keeping in the paper
Here is the part I want to be honest about, because it taught me more than the win did.
I expected fancier reward shaping to help. Composite rewards, and an MGPO variant that biases toward informative examples. Neither beat plain negative-CER GSPO. Not “beat it by a little,” did not beat it at all.
The reason is that low-resource ASR difficulty is bimodal, not smooth. The easy Spanish clips have a success probability near 1. The near-impossible Indigenous clips have a success probability near 0. In between there is only a thin frontier of clips where the model sometimes succeeds and sometimes fails, and the frontier is where learning signal lives. When I set a 0.5 difficulty threshold to find those frontier clips, roughly 23 of 50 batches had zero usable frontier clips. The fancier method was starving. It was engineered to exploit a gradient of difficulty that mostly is not there. Plain GSPO, which does not depend on that structure, kept learning.
The lesson generalizes: sophistication that assumes a smooth difficulty landscape is worse than useless on a bimodal one.
Two engineering war stories
The accidentally zero-shot preview. An earlier preview looked oddly weak, and it took a while to see why: it had been trained on one set of languages and evaluated on a different, non-overlapping set. The model was being graded, zero-shot, on languages it had literally never trained on. Lesson, now written down where I will see it: always confirm that your train and test label or language sets actually intersect before you believe any number.
The 30-second window. Evaluation clips run around 60 seconds, but Whisper’s attention window is 30. Naively you lose everything past the boundary. The fix is to force-align the transcript to the audio with MMS_FA, then cut on inter-word silences into chunks of at most 28 seconds, preserving the original orthography across the seams so the tone marks and glottal stops survive the split.
Where this sits
This is part of the broader low-resource line I publish openly on HuggingFace at Thermostatic, alongside the translation work. The models get replaced; the benchmark and the datasets are the durable contribution. MEXA exists so that the next person trying to make a speech model hear Nahuatl starts with a number they can trust, and a way to prove they did not cheat to get it.
Every top speech-recognition system on the planet can hear Spanish just fine. Speak Mexican Spanish into any of the big-name systems and it types out what you said with only about 14 percent of the words wrong, which is annoying but usable. Now point those exact same systems at Nahuatl, or any of the 22 other Indigenous languages of Mexico, and the wrong-word rate leaps to 99 percent or worse. At that point the system is not really listening. It is just guessing, and guessing in the wrong language entirely.
That gap, 14 percent versus 99 percent, is the whole project in a single number. And it is not a gap you can close by nudging a few settings. It exists because the practice material these systems need, hours of recorded speech paired with written transcripts, simply never existed for these 23 languages, which span at least six different language families. This is the story of closing part of that gap.
First, build a test that cannot be cheated
Before touching any AI, the first job was to build a fair exam, because in this kind of work a dishonest exam is where projects secretly rot. If you cannot tell real progress from random luck, every attempt looks like a triumph.
The exam is called MEXA, and it is built to be uncheatable in two ways. First, the answer key is kept private and never released, so nobody can peek. Second, there is a clever public “fingerprint” list: anyone can check whether their own practice recordings accidentally overlap with the secret exam, without ever seeing the exam itself. You learn only whether you have a collision, never what the answers are.
The scoring is also linguistically careful. Many of these languages use tone marks and little catches in the throat that completely change a word’s meaning. A lazy scorer would strip those marks away and quietly reward an AI for mangling them. This exam keeps them, so it grades the language people actually speak, not a flattened version of it.
Learning by practice and reward
The core idea borrows a training method usually reserved for chatbots: reinforcement learning. Here is the everyday version. Instead of forcing the AI to memorize a fixed answer sheet, you let it take a guess, then hand it a reward based on how good the guess was, then let it try again and again, steering toward whatever earns the biggest reward. It is learning by trial and feedback, the way you would learn a sport by practicing and being told how each attempt went, rather than by cramming from a book.
The training happened in two stages. The first stage was ordinary practice: show the AI recordings paired with correct transcripts until it stops guessing in Spanish and starts producing plausible-looking words in the right language. The second stage was the reinforcement-learning stage, using a method called GSPO, and this is where the interesting part lives.
The reward is beautifully simple
The reward the AI chases is just this: how many letters did it get right, compared to the correct transcript? That is the entire reward. There is no complicated human judge, no vague sense of “quality.” The correct transcript is ground truth, so the reward is exact and honest, which sidesteps a whole family of ways AIs learn to cheat their graders.
This simple reward had a lovely side effect. The worst failure of these systems on hard audio is getting stuck in a loop, repeating the same phrase over and over like a scratched record. Nobody had to tell the AI to stop looping. Under a letter-counting reward, every repeated letter counts as a mistake and shreds the score, so the AI discovers on its own that looping is expensive and quits doing it.
Why not just use an off-the-shelf tool?
The obvious shortcut was to grab an existing reinforcement-learning toolkit and be done by lunch. It did not work, for a deep reason. Every ready-made tool of this kind assumes the AI reads text and writes text. But this AI’s input is not text at all, it is sound. So the entire training machine had to be rebuilt by hand to feed the AI audio, let it produce several candidate transcripts per clip, and nudge it toward the better ones.
What happened, including the failure worth keeping
The reinforcement stage improved the results by a small but real margin, and crucially it dropped the worst looping language’s error rate from about 100 to about 76. That improvement was enough to take the number one spot on the public scoreboard, and it beat the next-best system on letter accuracy by a huge margin.
The most instructive part was a failure. The researcher expected that a fancier, cleverer reward would help even more. It did not, not even a little. The reason is honest: these languages come in two extremes, the easy Spanish clips the AI almost always gets right, and the near-impossible Indigenous clips it almost always gets wrong, with almost nothing in the middle. The fancy method was built to exploit a gentle gradient of difficulty that mostly did not exist. The plain reward kept learning precisely because it did not rely on that missing middle ground. Sometimes the sophisticated tool is worse than the simple one, and knowing why is the real prize.