Last Translation Benchmark
a crowdsourced, peer-reviewed set of hard-to-translate texts, images, audio and video, each graded by pass/fail verification rules
Vilém Zouhar, Niyati Bafna, Mukund Choudhary, Maike Züfle, Sara Rajaee, Pinzhen Chen, Jannis Vamvas, Sara Papi, Ona de Gibert, Bhavitvya Malik, Eliya Habba, Orfeas Menis Mastromichalakis, Patrícia Schmidtová, Michelle Wastl, Sheriff Issaka, Leshem Choshen, Stella Biderman, Antonis Anastasopoulos, Jan Niehues, Rico Sennrich, Mrinmaya Sachan, Ondřej Bojar, Kenton Murray, Jörg Tiedemann, Alham Fikri Aji, Philipp Koehn, Christof Monz, Alexandra Birch, Sowmya Vajjala, Chalamalasetti Kranti, Cristina España-Bonet, Nobin Sarwar, David Kaczér, Shunta Asano, Malik Marmonier, Daban Q. Jaff, Vaisakhi Mishra, Hend Al- Khalifa, Gabriele Sarti, Sourajit Saha, Nils Rehlinger, Juan Daniel Cuervo Villa, Jonathan Tonglet, Saugata Purkayastha, Dominik Macháček, J. Ramanujam, Heejin Do, Zuzana Nadova, Fred Philippy, Fabian Retkowski, Maria Lymperaiou, Silvia Casola, Hanna Yukhymenko, Shubhashis Roy Dipta, Sangwon Ryu, Andrés Jerez, Ron Keinan, Shuaib Shuaib Yusuf, Avantica Vempati, Maria Carmen Staiano, Sukannya Purkayastha, Adrian Cosma, Vitalii Babenko, Erivan Inan, Aviral Nigam, Wafa Aissa, Fatima Haouari, VENKATA PRASANTH KUMAR GUMMADI, Mehdi Jafarzadeh, Valentin Scourneau, Lukas Edman, Kaiser Sun, Shaomu Tan, Mohammad Gholizadeh, Johannes-Rudolf David, Dipankar Srirag, Javier García Gilabert, Ruta Binkyte, Manar Ali, Ana-Maria Bucur, Sabry E. Farrag, Youssef Saber, Yihong Liu, Jean Maillard, Cojocaru Nicoleta, Xiaochuang Yuan, Sina Ahmadi, Philipp Mondorf, Kaustubh Dhole, Roman Wixinger, Shenbin Qian, Manuel Tuor, Sergey Troshin, Jonathan Yahav, Fida Mohammad Thoker, Amir Rezapour, Lance Calvin Lim Gamboa, Manon Reusens, Kätriin Kukk, Koel Dutta Chowdhury · 2026
In one sentence
The Last Translation Benchmark is a live, crowdsourced collection of peer-reviewed examples that break leading machine translation models, each paired with handwritten pass/fail verification rules that an LLM checks, giving a reproducible, interpretable score instead of an opaque quality metric.
Abstract
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.
Questions this paper answers
- Is there a benchmark of sentences that today's best machine translation systems still get wrong?
- What challenge set targets failure modes of state-of-the-art MT now that standard test sets are saturating?
- How do I find hard, human-written test cases to stress-test a translation model?
- Which benchmark should I use if WMT-style test sets no longer separate my strong translation models?
- The Last Translation Benchmark (LTB) is a live machine translation challenge set of human-authored examples that most leading translation models fail on. Contributors keep adding examples, with tagged releases such as LTBv1.
Holds for: As of LTBv1, the release of contributions accepted before September 1st 2026; one of several community-driven efforts to collect model-breaking examples, applied here to translation.
- On LTBv1-eval, the best model, Gemini 3.1 Pro, passes all verification rules on only 37.4% to 43.8% of examples depending on the verifier LLM. GPT-5.6 Sol is second at 28.1% to 36.1%.
Holds for: 911 text-only LTBv1-eval examples in blind mode, scored by 6 verifier LLMs; examples a model cannot translate for lack of language support are excluded.
- How good are the best AI translators on really hard translation examples?
- What verifier pass rate do frontier LLMs such as Gemini 3.1 Pro and GPT-5.6 reach on LTBv1-eval?
- Can I trust Gemini or GPT to handle tricky idioms, puns and cultural references in translation?
- On LTBv1-eval, the best model, Gemini 3.1 Pro, passes all verification rules on only 37.4% to 43.8% of examples depending on the verifier LLM. GPT-5.6 Sol is second at 28.1% to 36.1%.
Holds for: 911 text-only LTBv1-eval examples in blind mode, scored by 6 verifier LLMs; examples a model cannot translate for lack of language support are excluded.
- The difficulty of the Last Translation Benchmark is not limited to the models examples were filtered against. 19 of the 29 evaluated models were never shown to contributors, and GPT-5.6 Luna among them passes only 15.4% to 22.9% of LTBv1-eval.
Holds for: Up to 10 models were shown on the contributor platform; the 29-model evaluation uses the same 6 verifiers on LTBv1-eval.
- Is a translation benchmark built by breaking a few models only hard for those same models?
- Does adversarial filtering against the platform models inflate difficulty only for those models, or does it transfer to held-out systems?
- If my translation model was never used to filter the Last Translation Benchmark, will it still find the examples hard?
- The difficulty of the Last Translation Benchmark is not limited to the models examples were filtered against. 19 of the 29 evaluated models were never shown to contributors, and GPT-5.6 Luna among them passes only 15.4% to 22.9% of LTBv1-eval.
Holds for: Up to 10 models were shown on the contributor platform; the 29-model evaluation uses the same 6 verifiers on LTBv1-eval.
- How can you score a translation without a vague quality number?
- How do example-specific pass/fail verification rules compare with reference-based metrics and rubric LLM judges for MT evaluation?
- How do I evaluate translations so that I know exactly which error a model made?
- Should I replace COMET or LLM-as-a-judge scores with checklist-style verification rules for my translation evaluation?
- The Last Translation Benchmark scores translations with example-specific verification rules, where a translation succeeds only if an LLM verifier finds it passes every rule. The rules give the evaluator privileged knowledge of the failure the generator lacks.
Holds for: Rules are written by the contributor of each example and describe its targeted failure, not every aspect of translation quality; LTBv1 averages 1.8 rules per example.
- Generic LLM judges score Gemini 3.1 Pro translations on LTBv1-eval at 72.1 to 95.4 out of 100, the rubric's good or very good bands. Rule-based verifiers pass at most 43.8% of them.
Holds for: Judges are the same 6 LLMs prompted with a cESA-style 0 to 100 rubric in which 65 to 80 means good; the 95.4 comes from Gemini 3.1 Pro judging its own output.
- Rankings of translation models by verifier pass rate agree with human judgments at an average Kendall tau of 90.5, versus 34.9 for generic LLM judges and 16.2 for automatic metrics.
Holds for: Human rankings from the cESA study on a subset of models and examples; the same agreement values are reported whether annotators saw the rules or not.
- Do AI judges of translation quality miss obvious translation mistakes?
- How do generic LLM-as-a-judge scores compare with rule-based verifier pass rates on hard MT examples?
- Is an LLM judge rating my translations as good enough evidence that they are correct?
- Generic LLM judges score Gemini 3.1 Pro translations on LTBv1-eval at 72.1 to 95.4 out of 100, the rubric's good or very good bands. Rule-based verifiers pass at most 43.8% of them.
Holds for: Judges are the same 6 LLMs prompted with a cESA-style 0 to 100 rubric in which 65 to 80 means good; the 95.4 comes from Gemini 3.1 Pro judging its own output.
- Standard translation metrics do not register the gain from giving LLM translators the verification rules. 2 of the 5 automatic metrics score the rule-informed translations lower, while the verifier pass rate rises from 7.2% to 89.8%.
Holds for: MetricX 24, MetricX QE 24, Comet 22, Comet QE 22 and ChrF, averaged over LLM translators on LTBv1-eval with the human translation as reference where needed.
- Do AI translators do better if you tell them what mistake to avoid?
- How much does conditioning an LLM translator on gold versus self-generated verification rules improve verifier pass rate?
- How do I get an LLM to avoid a known pitfall when translating a tricky sentence?
- Should I have my LLM write its own checklist of translation pitfalls before translating?
- Showing LLM translators the human-written verification rules raises the verifier pass rate on the Last Translation Benchmark from 7.2% to 89.8%. Rules the LLMs generate for themselves raise it only to 12.9%.
Holds for: Averaged over LLM translators on LTBv1-eval, in the oracle setting where the rules are privileged information not available to a real translation system.
- Do scores like COMET and chrF notice when a translation fixes the hard part?
- Are MetricX, COMET and chrF sensitive to translations that satisfy targeted verification rules?
- Can I rely on COMET or MetricX to tell me my translation model fixed a specific error?
- Standard translation metrics do not register the gain from giving LLM translators the verification rules. 2 of the 5 automatic metrics score the rule-informed translations lower, while the verifier pass rate rises from 7.2% to 89.8%.
Holds for: MetricX 24, MetricX QE 24, Comet 22, Comet QE 22 and ChrF, averaged over LLM translators on LTBv1-eval with the human translation as reference where needed.
- Does it matter which AI model you use to grade translations?
- How stable are MT system rankings across evaluator LLMs for rule verification, LLM judges and neural metrics, measured by Kendall tau?
- If I switch my evaluator LLM, will my translation model rankings change?
- Model rankings from verification rules are stable across the choice of verifier LLM, with an average Kendall tau of 86.9, against 72.0 for generic LLM judges and 35.1 for automatic metrics.
Holds for: Pairwise ranking similarity within each evaluation approach over the models on LTBv1-eval, using the 6 verifier or judge LLMs and 5 metrics.
- Rankings of translation models by verifier pass rate agree with human judgments at an average Kendall tau of 90.5, versus 34.9 for generic LLM judges and 16.2 for automatic metrics.
Holds for: Human rankings from the cESA study on a subset of models and examples; the same agreement values are reported whether annotators saw the rules or not.
- Do AI models grade their own translations too kindly?
- How large is self-preference bias of LLM evaluators in MT, as a judge versus as a rule verifier?
- Is it a problem if I use the same LLM to translate and to evaluate the translations?
- LLMs used as generic translation judges favour their own translations far more than when they act as rule verifiers. Gemma 4 shows 27.6% self-bias as a judge but 8.9% as a verifier.
Holds for: 6 LLMs, self-bias being self-ranking minus the average ranking by other models; Gemini 3.1 Pro instead shows negative self-bias in both roles (-8.3% judge, -8.9% verifier).
- Do human translators still beat AI on hard translation examples according to human raters?
- How do cESA human scores of reference translations compare with frontier MT output on LTBv1, with and without verification rules?
- Are human translations still clearly better than Gemini 3.1 Pro on tricky inputs?
- In a human evaluation, annotators rate the contributors' reference translations above every model they scored, 90.8 versus 68.1 for Gemini 3.1 Pro once shown the rules. Without the rules the gap is smaller, 81.4 versus 77.0.
Holds for: 22 bilingual annotators, 317 examples across 19 language pairs, a subset of models; cESA protocol, rating each translation first without and then with the rules.
- How big is the Last Translation Benchmark and which languages does it cover?
- What are the size, language coverage, modality mix and English-centricity of LTBv1?
- Does the Last Translation Benchmark have enough examples in my language pair to be useful to me?
- LTBv1 contains 3456 accepted examples spanning 109 languages, contributed by 260 people from 177 institutions. 94% of examples are text, and 13% involve neither English source nor English target.
Holds for: Contributions collected May to September 1st 2026; 73% translate into English and 14% out of English; the evaluation subset LTBv1-eval has 911 text-only examples.
- What kinds of text are hardest for machine translation today?
- Which sources of translation difficulty, such as metaphor, polysemy and cultural knowledge, dominate a crowdsourced MT challenge set?
- How do I find which linguistic phenomena my translation model is weakest on?
- The most frequent difficulty labels in LTBv1 are metaphor (1086 examples), cultural artifact (986) and polysemy (948). Less-benchmarked challenges also appear, including internet cultural artifacts (152) and meta-reasoning (83).
Holds for: Multi-label tags on all LTBv1 examples, assigned by an LLM after 2 linguists built the taxonomy inductively on a subset; tags are not exhaustive.
- How cheap is it to crowdsource hard translation test examples with AI checking?
- What is the model-call cost per accepted example in a crowdsourced MT challenge set with LLM verification?
- Can I afford to build a similar hard translation test set for my own domain?
- Collecting one accepted Last Translation Benchmark example costs about $0.12 in model calls. A contributor makes on average 10 translation attempts before submitting a valid example.
Holds for: Model-call costs only, for up to 10 platform translations plus Gemini 3.1 Pro verification of each rule; LTBv1 averages 1.8 rules per example.
Claims and scope
- The Last Translation Benchmark (LTB) is a live machine translation challenge set of human-authored examples that most leading translation models fail on. Contributors keep adding examples, with tagged releases such as LTBv1.
Scope: As of LTBv1, the release of contributions accepted before September 1st 2026; one of several community-driven efforts to collect model-breaking examples, applied here to translation.
- The Last Translation Benchmark scores translations with example-specific verification rules, where a translation succeeds only if an LLM verifier finds it passes every rule. The rules give the evaluator privileged knowledge of the failure the generator lacks.
Scope: Rules are written by the contributor of each example and describe its targeted failure, not every aspect of translation quality; LTBv1 averages 1.8 rules per example.
- On LTBv1-eval, the best model, Gemini 3.1 Pro, passes all verification rules on only 37.4% to 43.8% of examples depending on the verifier LLM. GPT-5.6 Sol is second at 28.1% to 36.1%. (Section 3.2)
Scope: 911 text-only LTBv1-eval examples in blind mode, scored by 6 verifier LLMs; examples a model cannot translate for lack of language support are excluded.
- The difficulty of the Last Translation Benchmark is not limited to the models examples were filtered against. 19 of the 29 evaluated models were never shown to contributors, and GPT-5.6 Luna among them passes only 15.4% to 22.9% of LTBv1-eval. (Section 3.2)
Scope: Up to 10 models were shown on the contributor platform; the 29-model evaluation uses the same 6 verifiers on LTBv1-eval.
- Generic LLM judges score Gemini 3.1 Pro translations on LTBv1-eval at 72.1 to 95.4 out of 100, the rubric's good or very good bands. Rule-based verifiers pass at most 43.8% of them. (Section 3.2)
Scope: Judges are the same 6 LLMs prompted with a cESA-style 0 to 100 rubric in which 65 to 80 means good; the 95.4 comes from Gemini 3.1 Pro judging its own output.
- Showing LLM translators the human-written verification rules raises the verifier pass rate on the Last Translation Benchmark from 7.2% to 89.8%. Rules the LLMs generate for themselves raise it only to 12.9%. (Table 3)
Scope: Averaged over LLM translators on LTBv1-eval, in the oracle setting where the rules are privileged information not available to a real translation system.
- Standard translation metrics do not register the gain from giving LLM translators the verification rules. 2 of the 5 automatic metrics score the rule-informed translations lower, while the verifier pass rate rises from 7.2% to 89.8%. (Table 3)
Scope: MetricX 24, MetricX QE 24, Comet 22, Comet QE 22 and ChrF, averaged over LLM translators on LTBv1-eval with the human translation as reference where needed.
- Model rankings from verification rules are stable across the choice of verifier LLM, with an average Kendall tau of 86.9, against 72.0 for generic LLM judges and 35.1 for automatic metrics. (Table 4)
Scope: Pairwise ranking similarity within each evaluation approach over the models on LTBv1-eval, using the 6 verifier or judge LLMs and 5 metrics.
- Rankings of translation models by verifier pass rate agree with human judgments at an average Kendall tau of 90.5, versus 34.9 for generic LLM judges and 16.2 for automatic metrics. (Table 4)
Scope: Human rankings from the cESA study on a subset of models and examples; the same agreement values are reported whether annotators saw the rules or not.
- LLMs used as generic translation judges favour their own translations far more than when they act as rule verifiers. Gemma 4 shows 27.6% self-bias as a judge but 8.9% as a verifier. (Table 5)
Scope: 6 LLMs, self-bias being self-ranking minus the average ranking by other models; Gemini 3.1 Pro instead shows negative self-bias in both roles (-8.3% judge, -8.9% verifier).
- In a human evaluation, annotators rate the contributors' reference translations above every model they scored, 90.8 versus 68.1 for Gemini 3.1 Pro once shown the rules. Without the rules the gap is smaller, 81.4 versus 77.0. (Section 3.2)
Scope: 22 bilingual annotators, 317 examples across 19 language pairs, a subset of models; cESA protocol, rating each translation first without and then with the rules.
- LTBv1 contains 3456 accepted examples spanning 109 languages, contributed by 260 people from 177 institutions. 94% of examples are text, and 13% involve neither English source nor English target. (Section 3; Table 1)
Scope: Contributions collected May to September 1st 2026; 73% translate into English and 14% out of English; the evaluation subset LTBv1-eval has 911 text-only examples.
- The most frequent difficulty labels in LTBv1 are metaphor (1086 examples), cultural artifact (986) and polysemy (948). Less-benchmarked challenges also appear, including internet cultural artifacts (152) and meta-reasoning (83). (Table 6; Section 3.4)
Scope: Multi-label tags on all LTBv1 examples, assigned by an LLM after 2 linguists built the taxonomy inductively on a subset; tags are not exhaustive.
- Collecting one accepted Last Translation Benchmark example costs about $0.12 in model calls. A contributor makes on average 10 translation attempts before submitting a valid example. (Section 2)
Scope: Model-call costs only, for up to 10 platform translations plus Gemini 3.1 Pro verification of each rule; LTBv1 averages 1.8 rules per example.
Common misreadings
- A low pass rate on the Last Translation Benchmark does not estimate typical translation quality: examples are selected to break leading models and rules target one failure each, so the benchmark is a stress test, not a measure of average user experience.
- The verifier pass rate is an LLM checking human-written rules, not a human judgment of every translation; its agreement with humans is shown on a 317-example human study, not on the full benchmark.
- The human reference translations pass the verification rules by construction, because a submission is accepted only if its reference passes; the separate human evaluation is what supports their superiority.
- The 89.8% pass rate with human rules in the prompt is an oracle result using privileged information, not a translation setup a deployed system can use; realistic comparisons use the blind mode.
- The name "Last Translation Benchmark" is hyperbole, according to its own contributor FAQ, and not a claim that machine translation evaluation is finished; the benchmark is live and gets new releases.
- The Last Translation Benchmark is intended for evaluation, not training, except for controlled research studies.
Terminology in this paper
- verification rule
- A short, English, pass/fail criterion written for one translation example that states the specific failure a correct translation must avoid, such as a required word sense or gender.
- verifier pass rate
- The percentage of examples for which a translation passes all of that example's verification rules, as judged by an LLM verifier.
- LTBv1
- The first tagged release of the Last Translation Benchmark, containing the 3456 examples accepted before September 1st 2026.
- LTBv1-eval
- A 911-example, text-only subset of LTBv1 selected for evaluation by difficulty, output diversity and balance across language pairs.
- blind mode
- A Last Translation Benchmark leaderboard setting in which the translation model sees only the input, used to compare realistic translation systems.
- oracle mode
- A Last Translation Benchmark leaderboard setting in which the translation model also sees privileged information such as the verification rules or the human translation.
- model blockers
- Taxonomy labels for target-side failures, such as refusals, irrelevant output, incomplete output, instruction injection and tokenization errors, that hide the source-side difficulty of an example.
How to cite
@misc{zouhar2026last,
title = {Last Translation Benchmark},
author = {Vilém Zouhar and Niyati Bafna and Mukund Choudhary and Maike Züfle and Sara Rajaee and Pinzhen Chen and Jannis Vamvas and Sara Papi and Ona de Gibert and Bhavitvya Malik and Eliya Habba and Orfeas Menis Mastromichalakis and Patrícia Schmidtová and Michelle Wastl and Sheriff Issaka and Leshem Choshen and Stella Biderman and Antonis Anastasopoulos and Jan Niehues and Rico Sennrich and Mrinmaya Sachan and Ondřej Bojar and Kenton Murray and Jörg Tiedemann and Alham Fikri Aji and Philipp Koehn and Christof Monz and Alexandra Birch and Sowmya Vajjala and Chalamalasetti Kranti and Cristina España-Bonet and Nobin Sarwar and David Kaczér and Shunta Asano and Malik Marmonier and Daban Q. Jaff and Vaisakhi Mishra and Hend Al- Khalifa and Gabriele Sarti and Sourajit Saha and Nils Rehlinger and Juan Daniel Cuervo Villa and Jonathan Tonglet and Saugata Purkayastha and Dominik Macháček and J. Ramanujam and Heejin Do and Zuzana Nadova and Fred Philippy and Fabian Retkowski and Maria Lymperaiou and Silvia Casola and Hanna Yukhymenko and Shubhashis Roy Dipta and Sangwon Ryu and Andrés Jerez and Ron Keinan and Shuaib Shuaib Yusuf and Avantica Vempati and Maria Carmen Staiano and Sukannya Purkayastha and Adrian Cosma and Vitalii Babenko and Erivan Inan and Aviral Nigam and Wafa Aissa and Fatima Haouari and VENKATA PRASANTH KUMAR GUMMADI and Mehdi Jafarzadeh and Valentin Scourneau and Lukas Edman and Kaiser Sun and Shaomu Tan and Mohammad Gholizadeh and Johannes-Rudolf David and Dipankar Srirag and Javier García Gilabert and Ruta Binkyte and Manar Ali and Ana-Maria Bucur and Sabry E. Farrag and Youssef Saber and Yihong Liu and Jean Maillard and Cojocaru Nicoleta and Xiaochuang Yuan and Sina Ahmadi and Philipp Mondorf and Kaustubh Dhole and Roman Wixinger and Shenbin Qian and Manuel Tuor and Sergey Troshin and Jonathan Yahav and Fida Mohammad Thoker and Amir Rezapour and Lance Calvin Lim Gamboa and Manon Reusens and Kätriin Kukk and Koel Dutta Chowdhury},
year = {2026},
doi = {10.48550/arxiv.2609.04173},
}
References
See the full reference list in the paper.