Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

a 42-language MMLU test set with human-verified translations and questions labelled as culturally sensitive or culturally agnostic

Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, Andre Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker · ACL 2025 · 2025

In one sentence

Global-MMLU re-annotates MMLU for the cultural, geographic and dialect knowledge its questions require and releases a 42-language test set with human-verified translations, so multilingual results can be reported separately on culturally-sensitive and culturally-agnostic subsets.

Abstract

Reliable multilingual evaluation is difficult, and culturally appropriate evaluation is even harder to achieve.A common practice to fill this gap is to machine-translate English evaluation sets. However, translation introduces language bias and carries over cultural and regional assumptions from the original questions {--} often testing knowledge irrelevant to the target audience. In this work, we highlight the extent and impact of these biases and present a multilingual evaluation framework that aims to mitigate them through improved translations and annotation practices.Through a large-scale study involving professional and community translators and annotators, we show that state-of-the-art models excel primarily by learning Western-centric concepts. Notably, we find that model rankings on the full MMLU change when evaluated on a subset of questions explicitly marked as culturally sensitive.We release Global MMLU, a multilingual extension of MMLU across 42 languages, featuring improved translation quality, expanded language coverage, and designated subsets labeled as culturally sensitive and culturally agnostic to enable a more comprehensive and equitable benchmark for evaluating language models across diverse linguistic and cultural contexts.

Questions this paper answers

how many questions on the MMLU exam benchmark need knowledge of a particular culture or region to answer?
what proportion of MMLU items are culturally sensitive, and which cultures do the culture-dependent items presuppose?
how do I find out whether an English multiple-choice benchmark I am reporting on is loaded with culture-specific questions?
should I worry that MMLU scores reflect Western cultural knowledge rather than general ability?
28% of MMLU questions require culturally sensitive knowledge, meaning cultural, geographic or dialect-specific knowledge, to be answered correctly. Geographic knowledge accounts for 54.7% of those questions and cultural knowledge for 32.7%.
Holds for: Based on a uniform sample of 2,850 English MMLU questions (50 per each of 57 subjects), labelled by majority vote among at least 3 of 200 professional and community annotators.
Among MMLU questions tagged as needing cultural knowledge, 86.5% require Western cultural knowledge, while the next largest category, South Asian culture, accounts for only 4%.
Holds for: English MMLU sample of 2,850 questions; percentages are over samples with a single culture tag, excluding untagged and multi-tag samples.
when MMLU questions depend on geography, which countries and continents do they actually talk about?
what is the regional distribution of geography-dependent MMLU items across North America, Europe and the rest of the world?
how do I check whether the geographic content of an exam-style benchmark is skewed toward the United States?
can I treat MMLU as a globally representative knowledge test, or is its geography mostly US-based?
For MMLU questions requiring geographic knowledge, 84.9% concern North America or Europe, split as 64.5% North America and 20.4% Europe. Of the Western-culture questions, 73.9% specifically require knowledge about the United States.
Holds for: English MMLU sample of 2,850 questions; region and country shares computed over samples with a single region or country tag.
which school subjects in a translated multiple-choice knowledge test contain the most culture-specific questions?
how is cultural sensitivity distributed across MMLU subject categories such as Humanities, Social Sciences and STEM?
how do I pick MMLU subjects that will not be confounded by cultural or regional knowledge?
if I only evaluate on STEM subjects of MMLU, do I avoid cultural bias?
Cultural sensitivity in MMLU is concentrated in the humanities: 68% of Humanities questions are culturally sensitive, versus 30 of 950 STEM samples (3.15%). All World Religions and Moral Scenarios samples contain at least one cultural, regional or dialect reference.
Holds for: English MMLU annotated sample of 2,850 questions, 50 per subject, with subject categories following Hendrycks et al. and Medical and Business split out of 'Other'. 12 of the 57 subjects contained no culturally sensitive samples.
do model leaderboards look different if you drop the questions that need culture-specific knowledge?
how much do model rankings shift between culturally-sensitive and culturally-agnostic MMLU subsets across languages?
how do I tell whether culture-dependent questions are distorting the model comparison I am publishing?
should I report separate scores for culture-dependent and culture-neutral MMLU questions when ranking models?
Model rankings shift far more on the culturally-sensitive MMLU subset than on the culturally-agnostic one. Averaged across languages and 14 models, CA rankings show 3.4 rank changes and 3.7 position shifts against the annotated MMLU sample, CS rankings 5.7 and 7.3.
Holds for: 14 models from 9 families (including GPT-4o and Claude 3.5 Sonnet) evaluated 5-shot with lm-evaluation-harness; ranks measured against the uniform MMLU Annotated subsample, across the 42 Global-MMLU languages.
are test scores across languages more erratic for languages with little data online?
how does cross-language accuracy standard deviation on MMLU compare between high-resource and low-resource languages?
how do I judge how much to trust a multilingual benchmark score for a low-resource language?
can I rely on a single multilingual MMLU number for a low-resource language, or is the spread too wide?
Accuracy variability across languages grows sharply for lower-resource languages. The average standard deviation rises from 3.21 (CA) and 3.86 (CS) on high-resource languages to 6.37 and 6.78 on low-resource languages, increases of 98% and 75%.
Holds for: 14 evaluated models, 5-shot; low-resource languages include machine-translated data, so part of the variance may reflect translation quality rather than model capability.
are questions that need cultural knowledge actually harder for language models to answer?
is average accuracy higher on the culturally-sensitive or the culturally-agnostic MMLU subset, and what explains the gap?
how should I read a higher score on culture-dependent MMLU questions than on culture-neutral ones?
if my model scores better on the culturally-sensitive split, does that mean it handles cultural knowledge well?
Average model accuracy is higher on the culturally-sensitive MMLU subset than on the culturally-agnostic one: 54.8% versus 51.3% for small models and 66.8% versus 61.6% for large models. CS questions come mostly from Social Sciences and Humanities, while CA retains the harder STEM and Medical subjects.
Holds for: 14 models, 5-shot, averaged across languages; small models are Aya Expanse 8B, Gemma2 9B, SEA-LION v3 9B, Llama 3.1 8B, Mistral Nemo 12B, Qwen2.5 7B and large models are Llama 3.1 70B and Command R+.
do language models score differently on questions translated by a machine than on questions translated by people?
how do model accuracies on human-translated versus machine-translated culturally-sensitive MMLU differ for Yoruba and French?
how do I decide whether machine translation is adequate for building an MMLU evaluation set in a low-resource language?
can I evaluate my model on machine-translated MMLU for Yoruba, or do I need human translators?
Claude 3.5 Sonnet and GPT-4o score significantly higher on machine-translated than on human-translated culturally-sensitive Yoruba data, while models generally do better on human-translated data for high-resource French; Aya Expanse 32B is the only model consistent across both.
Holds for: Compares human- versus machine-translated CS subsets for exactly 3 languages -- French (high), Korean (mid), Yoruba (low resource).
Accuracy variability across languages grows sharply for lower-resource languages. The average standard deviation rises from 3.21 (CA) and 3.86 (CS) on high-resource languages to 6.37 and 6.78 on low-resource languages, increases of 98% and 75%.
Holds for: 14 evaluated models, 5-shot; low-resource languages include machine-translated data, so part of the variance may reflect translation quality rather than model capability.
which translation tool produces better results when turning an English exam benchmark into other languages?
how do Google Translate and GPT-3.5-Turbo compare on ChrF++ against professional MMLU translations across subjects and languages?
how do I choose a translation system for producing a multilingual version of an English benchmark?
should I use Google Translate or an LLM to translate my evaluation set into 40-odd languages?
Google Translate achieves higher ChrF++ scores than GPT-3.5-Turbo across all MMLU subject categories and with lower deviation across languages, measured against professional human translations from MMMLU.
Holds for: Restricted to languages overlapping between the two machine-translated sets and human-translated MMMLU; GPT-3.5-Turbo is the system used for the widely adopted 26-language translated MMLU.
how much of a machine-translated exam benchmark did human reviewers actually have to fix?
what post-edit rate did professional annotators and community contributors apply to the machine-translated MMLU samples?
how do I estimate the human post-editing effort needed to clean up a machine-translated benchmark?
if I hire annotators to post-edit a translated benchmark, what share of samples should I budget for edits?
Human review changed a substantial share of the machine-translated MMLU: 7,565 edits were made, 36.9% of reviewed samples. Professional annotators edited 789 samples per language (38.5%) and community contributors 362 per language (17.7%).
Holds for: Edits cover 4 professionally annotated gold languages (Arabic, French, Hindi, Spanish) and 11 community-translated languages; differing edit rates reflect annotator time and resources, not translation quality across languages.
how many languages and questions are in the human-checked multilingual version of MMLU?
what is the language coverage and sample count of Global-MMLU, including its culturally-sensitive and culturally-agnostic annotated splits?
how do I find a multilingual knowledge benchmark that covers dozens of languages with human-verified translations?
is there a multilingual MMLU I can drop into my evaluation suite, and how big is it?
Global-MMLU covers all 14K MMLU samples in 42 languages, 589,764 samples in total, combining professional translations with post-edits, community translations and machine translation. It ships 792 English culturally-sensitive and 2,058 culturally-agnostic annotated questions, extended to the other 41 languages.
Holds for: Cultural sensitivity labels were assigned on the English source and propagated to translations, so they capture bias in the original questions rather than artefacts of any particular translation. Released under a permissive license.
what should I read first about evaluating language models across languages and cultures, not just translations?
which work established that machine-translating an English benchmark gives multilinguality without multiculturality?
how do I report multilingual evaluations separately for culture-dependent and culture-neutral questions?
which multilingual benchmark should I cite if I want to argue that translated evaluations miss cultural knowledge?
Global-MMLU is a reference point for the argument that machine-translating an English benchmark yields multilinguality without multiculturality, and provides the culturally-sensitive and culturally-agnostic subsets needed to report multilingual LLM evaluations separately.
Holds for: As of publication in 2025; concerns knowledge-style multiple-choice evaluation derived from English MMLU, covering 42 of the world's roughly 7,000 languages and no dialect variation.
does labelling which questions need cultural knowledge make a benchmark fair across cultures?
what are the limits of culturally-sensitive annotation on a translated benchmark for claims of cultural inclusivity?
how do I know whether filtering a translated benchmark by cultural sensitivity is enough for a fair cross-cultural evaluation?
if I use the culturally-agnostic split of a translated MMLU, can I claim my evaluation is culturally inclusive?
Identifying whether benchmark questions are culturally sensitive does not make a dataset culturally inclusive. Global-MMLU flags gaps in non-Western representation but still derives all its questions from English MMLU rather than from locally authored material.
Holds for: Stated as a limitation of the Global-MMLU release; achieving inclusion would require integrating culturally grounded knowledge sourced in-language, which the dataset does not do.

Claims and scope

Common misreadings

Terminology in this paper

Culturally-Sensitive (CS)
An MMLU question whose correct answer depends on cultural, geographic or dialect-specific knowledge, assigned by majority vote of annotators labelling the original English question.
Culturally-Agnostic (CA)
An MMLU question that requires none of cultural, geographic or dialect-specific knowledge to answer correctly, serving as a baseline subset for cross-language comparison.
MMLU Annotated (MA)
A uniform random subsample of 2,850 MMLU questions, 50 from each of the 57 subjects, annotated in English for cultural, geographic, dialect and temporal knowledge and used as the reference for model rankings.
transMMLU
Collective term for versions of MMLU produced by machine-translating the English dataset into other languages and used as-is for multilingual evaluation.
Gold Set
The 4 languages -- Arabic, French, Hindi and Spanish -- whose machine translations were reviewed and post-edited by compensated professional annotators.

How to cite

@inproceedings{singh2024global,
    title = "Global {MMLU}: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation",
    author = "Singh, Shivalika  and
      Romanou, Angelika  and
      Fourrier, Cl{\'e}mentine  and
      Adelani, David Ifeoluwa  and
      Ngui, Jian Gang  and
      Vila-Suero, Daniel  and
      Limkonchotiwat, Peerat  and
      Marchisio, Kelly  and
      Leong, Wei Qi  and
      Susanto, Yosephine  and
      Ng, Raymond  and
      Longpre, Shayne  and
      Ruder, Sebastian  and
      Ko, Wei-Yin  and
      Bosselut, Antoine  and
      Oh, Alice  and
      Martins, Andre  and
      Choshen, Leshem  and
      Ippolito, Daphne  and
      Ferrante, Enzo  and
      Fadaee, Marzieh  and
      Ermis, Beyza  and
      Hooker, Sara",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.919/",
    doi = "10.18653/v1/2025.acl-long.919",
    pages = "18761--18799",
    ISBN = "979-8-89176-251-0",
    abstract = "Reliable multilingual evaluation is difficult, and culturally appropriate evaluation is even harder to achieve.A common practice to fill this gap is to machine-translate English evaluation sets. However, translation introduces language bias and carries over cultural and regional assumptions from the original questions {--} often testing knowledge irrelevant to the target audience. In this work, we highlight the extent and impact of these biases and present a multilingual evaluation framework that aims to mitigate them through improved translations and annotation practices.Through a large-scale study involving professional and community translators and annotators, we show that state-of-the-art models excel primarily by learning Western-centric concepts. Notably, we find that model rankings on the full MMLU change when evaluated on a subset of questions explicitly marked as culturally sensitive.We release Global MMLU, a multilingual extension of MMLU across 42 languages, featuring improved translation quality, expanded language coverage, and designated subsets labeled as culturally sensitive and culturally agnostic to enable a more comprehensive and equitable benchmark for evaluating language models across diverse linguistic and cultural contexts."
}

References

See the full reference list in the paper.