Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
a hand-written commonsense reasoning benchmark for 141 language varieties, with culturally-specific and parallel splits
Tyler A. Chang, Catherine Arnett, Abdelrahman Eldesokey, Abdelrahman Sadallah, Abeer Kashar, Abolade Daud, Abosede Grace Olanihun, Adamu Labaran Mohammed, Adeyemi Praise, Adhikarinayum Meerajita Sharma, Aditi Gupta, Afitab Iyigun, Afonso Simplício, Ahmed Essouaied, Aicha Chorana, Akhil Eppa, Akintunde Oladipo, Akshay Ramesh, Aleksei Dorkin, Alfred Malengo Kondoro, Alham Fikri Aji, Ali Eren Çetintaş, Allan Hanbury, Alou Dembele, Alp Niksarli, Álvaro Arroyo, Amin Bajand, Amol Khanna, Ana Chkhaidze, Ana Condez, Andiswa Mkhonto, Andrew Hoblitzell, Andrew Tran, Angelos Poulis, Anirban Majumder, Anna Vacalopoulou, Annette Kuuipolani Kanahele Wong, Annika Simonsen, Anton Kovalev, Ashvanth. S, Ayodeji Joseph Lana, Barkin Kinay, Bashar Alhafni, Benedict Cibalinda Busole, Bernard Ghanem, Bharti Nathani, Biljana Stojanovska Duric, Bola Agbonile, Bragi Bergsson, Bruce Torres Fischer, Burak Tutar, Burcu Alakuş Çınar, Cade J. Kanoniakapueo Kane, Can Udomcharoenchaikit, Catherine Arnett, Chadi Helwe, Chaithra Reddy Nerella, Chen Cecilia Liu, Chiamaka Glory Nwokolo, Cristina España-Bonet, Cynthia Amol, DaeYeop Lee, Dana Arad, Daniil Dzenhaliou, Daria Pugacheva, Dasol Choi, Daud Abolade, David Liu, David Semedo, Deborah Popoola, Deividas Mataciunas, Delphine Nyaboke, Dhyuthy Krishna Kumar, Diogo Glória-Silva, Diogo Tavares, Divyanshu Goyal, DongGeon Lee, Ebele Nwamaka Anajemba, Egonu Ngozi Grace, Elena Mickel, Elena Tutubalina, Elias Herranen, Emile Anand, Emmanuel Habumuremyi, Emuobonuvie Maria Ajiboye, Eryawan Presma Yulianrifat, Esther Adenuga, Ewa Rudnicka, Faith Olabisi Itiola, Faran Taimoor Butt, Fathima Thekkekara, Fatima Haouari, Filbert Aurelian Tjiaranata, Firas Laakom, Francesca Grasso, Francesco Orabona, Francesco Periti, Gbenga Kayode Solomon, Gia Nghia Ngo, Gloria Udhehdhe-oze, Gonçalo Martins, Gopi Naga Sai Ram Challagolla, Guijin Son, Gulnaz Abdykadyrova, Hafsteinn Einarsson, Hai Hu, Hamidreza Saffari, Hamza Zaidi, Haopeng Zhang, Harethah Abu Shairah, Harry Vuong, Hele-Andra Kuulmets, Houda Bouamor, Hwanjo Yu, Iben Nyholm Debess, İbrahim Ethem Deveci, Ikhlasul Akmal Hanif, Ikhyun Cho, Inês Calvo, Inês Vieira, Isaac Manzi, Ismail Daud, Itay Itzhak, Iuliia, Alekseenko, Ivan Belashkin, Ivan Spada, Ivan Zhelyazkov, Jacob Brinton, Jafar Isbarov, Jaka Čibej, Jan Čuhel, Jan Kocoń, Jauza Akbar Krito, Jebish Purbey, Jennifer Mickel, Jennifer Za, Jenny Kunz, Jihae Jeong, Jimena Tena Dávalos, Jinu Lee, João Magalhães, John Yi, Jongin Kim, Joseph Chataignon, Joseph Marvin Imperial, Jubeerathan Thevakumar, Judith Land, Junchen Jiang, Jungwhan Kim, Kairit Sirts, Kamesh R, Kamesh V, Kanda Patrick Tshinu, Kätriin Kukk, Kaustubh Ponkshe, Kavsar Huseynova, Ke He, Kelly Buchanan, Kengatharaiyer Sarveswaran, Kerem Zaman, Khalil Mrini, Kian Kyars, Krister Kruusmaa, Kusum Chouhan, Lainitha Krishnakumar, Laura Castro Sánchez, Laura Porrino Moscoso, Leshem Choshen, Levent Sencan, Lilja Øvrelid, Lisa Alazraki, Lovina Ehimen-Ugbede, Luheerathan Thevakumar, Luxshan Thavarasa, Mahnoor Malik, Mamadou K. Keita, Mansi Jangid, Marco De Santis, Marcos García, Marek Suppa, Mariam D'Ciofalo, Marii Ojastu, Maryam Sikander, Mausami Narayan, Maximos Skandalis, Mehak Mehak, Mehmet İlteriş Bozkurt, Melaku Bayu Workie, Menan Velayuthan, Michael Leventhal, Michał Marcińczuk, Mirna Potočnjak, Mohammadamin Shafiei, Mridul Sharma, Mrityunjaya Indoria, Muhammad Ravi Shulthan Habibi, Murat Kolić, Nada Galant, Naphat Permpredanun, Narada Maugin, Nicholas Kluge Corrêa, Nikola Ljubešić, Nirmal Thomas, Nisansa de Silva, Nisheeth Joshi, Nitish Ponkshe, Nizar Habash, Nneoma C. Udeze, Noel Thomas, Noémi Ligeti-Nagy, Nouhoum Coulibaly, Nsengiyumva Faustin, Odunayo Kareemat Buliaminu, Odunayo Ogundepo, Oghojafor Godswill Fejiro, Ogundipe Blessing Funmilola, Okechukwu God'spraise, Olanrewaju Samuel, Olaoye Deborah Oluwaseun, Olasoji Akindejoye, Olga Popova, Olga Snissarenko, Onyinye Anulika Chiemezie, Orkun Kinay, Osman Tursun, Owoeye Tobiloba Moses, Oyelade Oluwafemi Joshua, Oyesanmi Fiyinfoluwa, Pablo Gamallo, Pablo Rodríguez Fernández, Palak Arora, Pedro Valente, Peter Rupnik, Philip Oghenesuowho Ekiugbo, Pramit Sahoo, Prokopis Prokopidis, Pua Niau-Puhipau, Quadri Yahya, Rachele Mignone, Raghav Singhal, Ram Mohan Rao Kadiyala, Raphael Merx, Rapheal Afolayan, Ratnavel Rajalakshmi, Rishav Ghosh, Romina Oji, Ron Kekeha Solis, Rui Guerra, Rushikesh Zawar, Sa'ad Nasir Bashir, Saeed Alzaabi, Sahil Sandeep, Sai Pavan Batchu, SaiSandeep Kantareddy, Salsabila Zahirah Pranida, Sam Buchanan, Samuel Rutunda, Sander Land, Sarah Sulollari, Sardar Ali, Saroj Sapkota, Saulius Tautvaisas, Sayambhu Sen, Sayantani Banerjee, Sebastien Diarra, SenthilNathan. M, Sewoong Lee, Shaan Shah, Shankar Venkitachalam, Sharifa Djurabaeva, Sharon Ibejih, Shivanya Shomir Dutta, Siddhant Gupta, Silvia Paniagua Suárez, Sina Ahmadi, Sivasuthan Sukumar, Siyuan Song, Snegha A., Sokratis Sofianopoulos, Sona Elza Simon, Sonja Benčina, Sophie Gvasalia, Sphurti Kirit More, Spyros Dragazis, Stephan P. Kaufhold, Suba. S, Sultan AlRashed, Surangika Ranathunga, Taiga Someya, Taja Kuzman Pungeršek, Tal Haklay, Tasi'u Jibril, Tatsuya Aoyama, Tea Abashidze, Terenz Jomar Dela Cruz, Terra Blevins, Themistoklis Nikas, Theresa Dora Idoko, Thu Mai Do, Tilek Chubakov, Tommaso Gargiani, Uma Rathore, Uni Johannesen, Uwuma Doris Ugwu, Vallerie Alexandra Putra, Vanya Bannihatti Kumar, Varsha Jeyarajalingam, Varvara Arzt, Vasudevan Nedumpozhimana, Viktoria Ondrejova, Viktoryia Horbik, Vishnu Vardhan Reddy Kummitha, Vuk Dinić, Walelign Tewabe Sewunetie, Winston Wu, Xiaojing Zhao, Yacouba Diarra, Yaniv Nikankin, Yash Mathur, Yixi Chen, Yiyuan Li, Yolanda Xavier, Yonatan Belinkov, Yusuf Ismail Abayomi, Zaid Alyafeai, Zhengyang Shan, Zhi Rui Tam, Zilu Tang, Zuzana Nadova, Baber Abbasi, Stella Biderman, David Stap, Duygu Ataman, Fabian Schmidt, Hila Gonen, Jiayi Wang, David Ifeoluwa Adelani · arXiv · 2025
In one sentence
Global PIQA is a commonsense reasoning benchmark for 141 language varieties, hand-written by over 350 researchers in over 65 countries, with a culturally-specific non-parallel split (100 examples per language) and a parallel split of 103 "culturally agnostic" questions translated into 131 language varieties.
Abstract
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we present Global PIQA, a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The 141 language varieties in Global PIQA cover five continents, 19 language families, and 24 writing systems. In the non-parallel split of Global PIQA, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements. In the parallel split, we translate more"culturally agnostic"commonsense reasoning questions into 131 language varieties, for direct cross-lingual comparisons. In both splits, all examples have been verified by native speakers of the languages. We find that state-of-the-art LLMs perform well on Global PIQA in aggregate, but they exhibit weaker performance in lower-resource languages (e.g. up to a 68% accuracy gap between languages in the parallel split). Global PIQA highlights that in many languages and cultures, everyday knowledge remains an area for improvement in LLMs, alongside more widely-discussed capabilities such as complex reasoning and expert knowledge. Beyond its uses for LLM evaluation, Global PIQA provides a glimpse into the wide diversity of cultures in which human language is embedded.
Questions this paper answers
- What benchmark should I use to test commonsense reasoning in low-resource languages?
- Is there a multilingual commonsense reasoning dataset that is not translated from English?
- Where should I start reading about culturally-specific LLM evaluation across many languages?
- Global PIQA departs from multilingual benchmarks such as XNLI, XCOPA, Belebele, MGSM and Global MMLU by writing its non-parallel split directly in each language instead of translating an English dataset. Examples translated from English PIQA are excluded from that split.
Holds for: The non-parallel split, covering 136 language varieties; the parallel split is itself translated from English by design, and a few non-parallel datasets were translated between related languages within the project.
- Global PIQA was built as a participatory benchmark: over 350 researchers from over 65 countries and over 180 affiliations wrote and validated examples in their own languages. Contributors were offered co-authorship rather than paid as external annotators.
Holds for: Participation was voluntary, recruited through NLP community channels such as the EleutherAI Discord, LINGUIST List, Masakhane and social media, between the June 2025 announcement and the September 15, 2025 deadline.
- Global PIQA covers 141 language varieties, spanning 118 unique ISO 639-3 language codes, 24 writing systems, and 19 top-level Glottolog language families. Of the languages, 70 are Indo-European and 11 Atlantic-Congo.
Holds for: Counts include dialect region codes; excluding them gives 129 ISO language-script combinations. Coverage is uneven: 36 South Asian languages against 1 Oceanian, and no indigenous American languages.
- How many languages does Global PIQA cover?
- Which language families and writing systems are represented in Global PIQA?
- How broad is the language coverage of a hand-written multilingual commonsense benchmark?
- Global PIQA covers 141 language varieties, spanning 118 unique ISO 639-3 language codes, 24 writing systems, and 19 top-level Glottolog language families. Of the languages, 70 are Indo-European and 11 Atlantic-Congo.
Holds for: Counts include dialect region codes; excluding them gives 129 ISO language-script combinations. Coverage is uneven: 36 South Asian languages against 1 Oceanian, and no indigenous American languages.
- How was Global PIQA constructed?
- Who wrote the examples in Global PIQA, and were they paid annotators?
- How do you organise a benchmark written by hundreds of native speakers?
- Global PIQA was built as a participatory benchmark: over 350 researchers from over 65 countries and over 180 affiliations wrote and validated examples in their own languages. Contributors were offered co-authorship rather than paid as external annotators.
Holds for: Participation was voluntary, recruited through NLP community channels such as the EleutherAI Discord, LINGUIST List, Masakhane and social media, between the June 2025 announcement and the September 15, 2025 deadline.
- Of the 146 author groups contributing datasets to the Global PIQA non-parallel split, 128 drafted their examples entirely manually without LLM help, and 141 reported making their datasets at least partially culturally-specific.
Holds for: Self-reported per-group method descriptions, non-parallel split only; 16 groups used LLMs for initial generation, with two reporting that they kept only 14.6% and 22.0% of generated examples.
- How much of Global PIQA is actually culturally specific?
- What fraction of examples reference local foods, customs or traditions?
- Were LLMs used to write the Global PIQA examples?
- In the official non-parallel split of Global PIQA, 59.9% of examples are annotated as culturally-specific, referencing local foods, holidays, folklore, traditions or region-varying norms. Only 4.1% were written with any help from LLMs.
Holds for: The official split up-samples culturally-specific and non-LLM examples, so the unsampled 29.1K-example pool is only 40.8% culturally-specific and 9.6% LLM-assisted.
- Of the 146 author groups contributing datasets to the Global PIQA non-parallel split, 128 drafted their examples entirely manually without LLM help, and 141 reported making their datasets at least partially culturally-specific.
Holds for: Self-reported per-group method descriptions, non-parallel split only; 16 groups used LLMs for initial generation, with two reporting that they kept only 14.6% and 22.0% of generated examples.
- How was data quality verified in a benchmark built by hundreds of volunteers?
- Were the Global PIQA examples checked by native speakers?
- Do the Global PIQA examples come with English translations?
- Every example in the Global PIQA non-parallel split was manually validated by at least one native speaker, 97.8% were validated multiple times by native speakers, and 92.6% carry human-corrected English translations.
Holds for: Secondary review, covering source-language verification and translation correction, was completed for 126 of the 136 non-parallel language varieties; the rest are marked [machine_translated] in the release.
- What is the difference between the parallel and non-parallel splits of Global PIQA?
- How can I compare LLM accuracy directly across languages on commonsense reasoning?
- How many examples are in the Global PIQA parallel split?
- The parallel split of Global PIQA contains 103 four-choice "culturally agnostic" commonsense questions in 131 language varieties. All were machine translated from English and then corrected or verified by a native speaker of each target language.
Holds for: Ekpeye has 101 examples because 2 cardinal-direction questions had no translatable terms; 6 of the original 109 English examples were dropped from all languages after correction revealed ambiguity.
- Global PIQA departs from multilingual benchmarks such as XNLI, XCOPA, Belebele, MGSM and Global MMLU by writing its non-parallel split directly in each language instead of translating an English dataset. Examples translated from English PIQA are excluded from that split.
Holds for: The non-parallel split, covering 136 language varieties; the parallel split is itself translated from English by design, and a few non-parallel datasets were translated between related languages within the project.
- How much do machine translations of commonsense questions need correcting for low-resource languages?
- Which languages required the heaviest edits to machine-translated benchmark examples?
- Is machine translation good enough for building multilingual benchmarks?
- Correcting the machine translations for the Global PIQA parallel split changed a mean of 24.9 characters per example, or 12.9% of characters. Per-language means reach 273.7 characters for Ekpeye, 209.0 for Idoma and 131.6 for Urhobo.
Holds for: Translations came from Gemini 2.5 Pro for the first 50 examples per language and Gemini 3.0 Flash for the rest, with prompts and candidate solutions translated separately.
- The parallel split of Global PIQA contains 103 four-choice "culturally agnostic" commonsense questions in 131 language varieties. All were machine translated from English and then corrected or verified by a native speaker of each target language.
Holds for: Ekpeye has 101 examples because 2 cardinal-direction questions had no translatable terms; 6 of the original 109 English examples were dropped from all languages after correction revealed ambiguity.
- How well do frontier LLMs do on multilingual commonsense reasoning?
- What accuracy do GPT-5.4, Claude and Gemini get on Global PIQA?
- Is Global PIQA already saturated by closed models?
- Some closed systems among GPT-5.4, Claude Sonnet 4.6 and Gemini 3.1 exceed 90% accuracy averaged across languages on both the parallel and non-parallel splits of Global PIQA.
Holds for: Generation-style prompting with thinking enabled (1024-token budget for Gemini and Claude, "medium" for GPT-5.4), 100 examples per language non-parallel and 103 parallel.
- Taking the best-performing LLM per language including closed systems, Global PIQA leaves 4 parallel-split and 8 non-parallel-split languages below 80% accuracy. Ekpeye sits at 33% parallel / 65% non-parallel and Idoma at 37% / 75%.
Holds for: Best-of-all-models-per-language figures over 100-103 examples per language, so sampling error is non-trivial; other low scorers include Burushaski at 59% and Meitei Manipuri at 63% non-parallel.
- What is the best open-weight model on multilingual commonsense reasoning?
- How big is the gap between open-weight models and proprietary systems on Global PIQA?
- Does scaling open-weight model size keep improving multilingual commonsense accuracy?
- Gemma 4 31B, the best open-weight model evaluated on Global PIQA, reaches 82.4% mean accuracy on the parallel split and 84.9% on the non-parallel split, below the closed-system skyline. Open-weight accuracy plateaus around 30-40B parameters.
Holds for: Open-weight models from 300M to 120B parameters with generation-style prompting, including models specialised for single languages or regions; the plateau is read against parameter count.
- How large is the accuracy gap between high- and low-resource languages on commonsense reasoning?
- Do LLMs perform worse on Sub-Saharan African languages than European ones?
- What accuracy disparity across regions does Global PIQA reveal?
- On the parallel split of Global PIQA, Gemma 4 31B averages 88.1% accuracy for European languages but only 60.5% for Sub-Saharan African languages, and 91.0% for high-resource against 75.0% for low-resource languages.
Holds for: Parallel split only, so the gap is not attributable to cultural content; resource levels follow the Joshi et al. (2020) taxonomy and regions follow the paper's own grouping.
- Taking the best-performing LLM per language including closed systems, Global PIQA leaves 4 parallel-split and 8 non-parallel-split languages below 80% accuracy. Ekpeye sits at 33% parallel / 65% non-parallel and Idoma at 37% / 75%.
Holds for: Best-of-all-models-per-language figures over 100-103 examples per language, so sampling error is non-trivial; other low scorers include Burushaski at 59% and Meitei Manipuri at 63% non-parallel.
- Which languages do LLMs handle worst on Global PIQA?
- Are there languages where even the best model scores under 80% on commonsense questions?
- How badly do models do on Ekpeye and Idoma?
- Taking the best-performing LLM per language including closed systems, Global PIQA leaves 4 parallel-split and 8 non-parallel-split languages below 80% accuracy. Ekpeye sits at 33% parallel / 65% non-parallel and Idoma at 37% / 75%.
Holds for: Best-of-all-models-per-language figures over 100-103 examples per language, so sampling error is non-trivial; other low scorers include Burushaski at 59% and Meitei Manipuri at 63% non-parallel.
- Can a benchmark separate a model's cultural knowledge from its linguistic ability in a language?
- Which languages show weaker cultural knowledge than linguistic competence in LLMs?
- What does the drop from the parallel to the non-parallel split of Global PIQA mean?
- Lingala and Plateau Malagasy show the largest parallel-to-non-parallel accuracy drops in Global PIQA for the best-performing models, at 20 and 19 points. The paper reads this as weaker cultural knowledge than linguistic ability.
Holds for: Best-model-per-language comparison across the two splits; the non-parallel split varies qualitatively in difficulty across languages, so split-to-split drops are suggestive rather than controlled.
- The parallel split of Global PIQA contains 103 four-choice "culturally agnostic" commonsense questions in 131 language varieties. All were machine translated from English and then corrected or verified by a native speaker of each target language.
Holds for: Ekpeye has 101 examples because 2 cardinal-direction questions had no translatable terms; 6 of the original 109 English examples were dropped from all languages after correction revealed ambiguity.
- Is everyday commonsense still a weakness of LLMs, or only expert reasoning?
- What does Global PIQA claim about where multilingual LLMs still fail?
- Why evaluate commonsense knowledge rather than complex reasoning across languages?
- Global PIQA argues that everyday commonsense knowledge, not only complex reasoning and expert knowledge, remains an area for improvement in LLMs for many languages and cultures.
Holds for: Based on 100 examples per language in one task format (prompt plus candidate solutions), for models evaluated in 2025-2026.
- On the parallel split of Global PIQA, Gemma 4 31B averages 88.1% accuracy for European languages but only 60.5% for Sub-Saharan African languages, and 91.0% for high-resource against 75.0% for low-resource languages.
Holds for: Parallel split only, so the gap is not attributable to cultural content; resource levels follow the Joshi et al. (2020) taxonomy and regions follow the paper's own grouping.
Claims and scope
- Global PIQA covers 141 language varieties, spanning 118 unique ISO 639-3 language codes, 24 writing systems, and 19 top-level Glottolog language families. Of the languages, 70 are Indo-European and 11 Atlantic-Congo. (Table 2 and Appendix B)
Scope: Counts include dialect region codes; excluding them gives 129 ISO language-script combinations. Coverage is uneven: 36 South Asian languages against 1 Oceanian, and no indigenous American languages.
- Global PIQA was built as a participatory benchmark: over 350 researchers from over 65 countries and over 180 affiliations wrote and validated examples in their own languages. Contributors were offered co-authorship rather than paid as external annotators. (Section 3.1 and Appendix C)
Scope: Participation was voluntary, recruited through NLP community channels such as the EleutherAI Discord, LINGUIST List, Masakhane and social media, between the June 2025 announcement and the September 15, 2025 deadline.
- In the official non-parallel split of Global PIQA, 59.9% of examples are annotated as culturally-specific, referencing local foods, holidays, folklore, traditions or region-varying norms. Only 4.1% were written with any help from LLMs. (Section 3.3 and Appendix D.2)
Scope: The official split up-samples culturally-specific and non-LLM examples, so the unsampled 29.1K-example pool is only 40.8% culturally-specific and 9.6% LLM-assisted.
- Every example in the Global PIQA non-parallel split was manually validated by at least one native speaker, 97.8% were validated multiple times by native speakers, and 92.6% carry human-corrected English translations. (Section 3.4)
Scope: Secondary review, covering source-language verification and translation correction, was completed for 126 of the 136 non-parallel language varieties; the rest are marked [machine_translated] in the release.
- Of the 146 author groups contributing datasets to the Global PIQA non-parallel split, 128 drafted their examples entirely manually without LLM help, and 141 reported making their datasets at least partially culturally-specific. (Section 3.2)
Scope: Self-reported per-group method descriptions, non-parallel split only; 16 groups used LLMs for initial generation, with two reporting that they kept only 14.6% and 22.0% of generated examples.
- The parallel split of Global PIQA contains 103 four-choice "culturally agnostic" commonsense questions in 131 language varieties. All were machine translated from English and then corrected or verified by a native speaker of each target language. (Section 4 and Section 4.1)
Scope: Ekpeye has 101 examples because 2 cardinal-direction questions had no translatable terms; 6 of the original 109 English examples were dropped from all languages after correction revealed ambiguity.
- Correcting the machine translations for the Global PIQA parallel split changed a mean of 24.9 characters per example, or 12.9% of characters. Per-language means reach 273.7 characters for Ekpeye, 209.0 for Idoma and 131.6 for Urhobo. (Section 4)
Scope: Translations came from Gemini 2.5 Pro for the first 50 examples per language and Gemini 3.0 Flash for the rest, with prompts and candidate solutions translated separately.
- Some closed systems among GPT-5.4, Claude Sonnet 4.6 and Gemini 3.1 exceed 90% accuracy averaged across languages on both the parallel and non-parallel splits of Global PIQA. (Section 5.3 and Figure 2)
Scope: Generation-style prompting with thinking enabled (1024-token budget for Gemini and Claude, "medium" for GPT-5.4), 100 examples per language non-parallel and 103 parallel.
- Gemma 4 31B, the best open-weight model evaluated on Global PIQA, reaches 82.4% mean accuracy on the parallel split and 84.9% on the non-parallel split, below the closed-system skyline. Open-weight accuracy plateaus around 30-40B parameters. (Figure 2 and Section 5.3)
Scope: Open-weight models from 300M to 120B parameters with generation-style prompting, including models specialised for single languages or regions; the plateau is read against parameter count.
- On the parallel split of Global PIQA, Gemma 4 31B averages 88.1% accuracy for European languages but only 60.5% for Sub-Saharan African languages, and 91.0% for high-resource against 75.0% for low-resource languages. (Figure 3 and Section 5.3)
Scope: Parallel split only, so the gap is not attributable to cultural content; resource levels follow the Joshi et al. (2020) taxonomy and regions follow the paper's own grouping.
- Taking the best-performing LLM per language including closed systems, Global PIQA leaves 4 parallel-split and 8 non-parallel-split languages below 80% accuracy. Ekpeye sits at 33% parallel / 65% non-parallel and Idoma at 37% / 75%. (Section 5.3 and Table 1)
Scope: Best-of-all-models-per-language figures over 100-103 examples per language, so sampling error is non-trivial; other low scorers include Burushaski at 59% and Meitei Manipuri at 63% non-parallel.
- Lingala and Plateau Malagasy show the largest parallel-to-non-parallel accuracy drops in Global PIQA for the best-performing models, at 20 and 19 points. The paper reads this as weaker cultural knowledge than linguistic ability. (Section 5.3)
Scope: Best-model-per-language comparison across the two splits; the non-parallel split varies qualitatively in difficulty across languages, so split-to-split drops are suggestive rather than controlled.
- Global PIQA argues that everyday commonsense knowledge, not only complex reasoning and expert knowledge, remains an area for improvement in LLMs for many languages and cultures. (Section 6)
Scope: Based on 100 examples per language in one task format (prompt plus candidate solutions), for models evaluated in 2025-2026.
- Global PIQA departs from multilingual benchmarks such as XNLI, XCOPA, Belebele, MGSM and Global MMLU by writing its non-parallel split directly in each language instead of translating an English dataset. Examples translated from English PIQA are excluded from that split. (Section 1 and Section 3.2)
Scope: The non-parallel split, covering 136 language varieties; the parallel split is itself translated from English by design, and a few non-parallel datasets were translated between related languages within the project.
Common misreadings
- Top closed systems exceeding 90% average accuracy does not mean Global PIQA is saturated: the benchmark still separates closed systems from open models, separates open models from each other, and leaves 8 non-parallel-split languages below 80% for the best model available.
- Global PIQA is not a translation of English PIQA. Examples in the non-parallel split were written directly in each language, and translated English PIQA examples were explicitly excluded from that split.
- The 141 language varieties are not 141 distinct ISO 639-3 languages: the count includes script variants and dialect region codes, and reduces to 129 language-script combinations and 118 ISO 639-3 codes.
- An example annotated as culturally-specific in Global PIQA does not always require the referenced cultural knowledge to answer; in many cases the culturally-specific element is only mentioned and the correct answer can be inferred from the rest of the context.
- The low accuracies for languages such as Ekpeye and Idoma on the parallel split are not explained by culturally unfamiliar content, because the parallel split was written to be culturally agnostic and translated from the same English source.
- Global PIQA's culturally-specific examples are snapshots from its individual authors and are not claimed to be representative of entire cultures; cultural stereotypes may still be present.
- "More languages is better" is not the Global PIQA authors' position: the paper states that researchers should work with communities to decide if and how their languages are included.
Terminology in this paper
- non-parallel split
- The portion of Global PIQA whose examples are written directly in each language by native speakers rather than translated, so examples differ across languages and can be culturally specific; 100 examples per language for 136 language varieties.
- parallel split
- The portion of Global PIQA consisting of the same 103 four-choice commonsense questions translated from English into 131 language varieties, enabling direct cross-lingual accuracy comparisons.
- culturally-specific example
- In Global PIQA, an example that uses words that do not translate well into English (e.g. local dishes or brands), describes specific holidays, folklore, traditions or sayings, or whose correct solution likely varies by region.
- culturally agnostic
- Written so as to minimise references to local foods, customs or traditions, so that the same question can be translated into a large number of languages and remain valid.
- cloze evaluation
- Scoring a candidate solution by the language model's log-probability of the solution given the prompt, normalised by the solution's length in bytes, used for pretrained-only models.
- generation evaluation
- Prompting an instruction-tuned model with the question and candidate solutions, sampling up to 2048 tokens, and scoring the response by string matching.
- English byte equivalents
- Text length in UTF-8 bytes divided by the language's byte premium, i.e. the estimated extra bytes that language needs relative to content-matched English, used to compare solution lengths across languages fairly.
How to cite
@misc{chang2025globalpiqaevaluatingphysical,
title={Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures},
author={Tyler A. Chang and Catherine Arnett and Abdelrahman Eldesokey and Abdelrahman Sadallah and Abeer Kashar and Abolade Daud and Abosede Grace Olanihun and Adamu Labaran Mohammed and Adeyemi Praise and Adhikarinayum Meerajita Sharma and Aditi Gupta and Afitab Iyigun and Afonso Simplício and Ahmed Essouaied and Aicha Chorana and Akhil Eppa and Akintunde Oladipo and Akshay Ramesh and Aleksei Dorkin and Alfred Malengo Kondoro and Alham Fikri Aji and Ali Eren Çetintaş and Allan Hanbury and Alou Dembele and Alp Niksarli and Álvaro Arroyo and Amin Bajand and Amol Khanna and Ana Chkhaidze and Ana Condez and Andiswa Mkhonto and Andrew Hoblitzell and Andrew Tran and Angelos Poulis and Anirban Majumder and Anna Vacalopoulou and Annette Kuuipolani Kanahele Wong and Annika Simonsen and Anton Kovalev and Ashvanth. S and Ayodeji Joseph Lana and Barkin Kinay and Bashar Alhafni and Benedict Cibalinda Busole and Bernard Ghanem and Bharti Nathani and Biljana Stojanovska Duric and Bola Agbonile and Bragi Bergsson and Bruce Torres Fischer and Burak Tutar and Burcu Alakuş Çınar and Cade J. Kanoniakapueo Kane and Can Udomcharoenchaikit and Catherine Arnett and Chadi Helwe and Chaithra Reddy Nerella and Chen Cecilia Liu and Chiamaka Glory Nwokolo and Cristina España-Bonet and Cynthia Amol and DaeYeop Lee and Dana Arad and Daniil Dzenhaliou and Daria Pugacheva and Dasol Choi and Daud Abolade and David Liu and David Semedo and Deborah Popoola and Deividas Mataciunas and Delphine Nyaboke and Dhyuthy Krishna Kumar and Diogo Glória-Silva and Diogo Tavares and Divyanshu Goyal and DongGeon Lee and Ebele Nwamaka Anajemba and Egonu Ngozi Grace and Elena Mickel and Elena Tutubalina and Elias Herranen and Emile Anand and Emmanuel Habumuremyi and Emuobonuvie Maria Ajiboye and Eryawan Presma Yulianrifat and Esther Adenuga and Ewa Rudnicka and Faith Olabisi Itiola and Faran Taimoor Butt and Fathima Thekkekara and Fatima Haouari and Filbert Aurelian Tjiaranata and Firas Laakom and Francesca Grasso and Francesco Orabona and Francesco Periti and Gbenga Kayode Solomon and Gia Nghia Ngo and Gloria Udhehdhe-oze and Gonçalo Martins and Gopi Naga Sai Ram Challagolla and Guijin Son and Gulnaz Abdykadyrova and Hafsteinn Einarsson and Hai Hu and Hamidreza Saffari and Hamza Zaidi and Haopeng Zhang and Harethah Abu Shairah and Harry Vuong and Hele-Andra Kuulmets and Houda Bouamor and Hwanjo Yu and Iben Nyholm Debess and İbrahim Ethem Deveci and Ikhlasul Akmal Hanif and Ikhyun Cho and Inês Calvo and Inês Vieira and Isaac Manzi and Ismail Daud and Itay Itzhak and Iuliia and Alekseenko and Ivan Belashkin and Ivan Spada and Ivan Zhelyazkov and Jacob Brinton and Jafar Isbarov and Jaka Čibej and Jan Čuhel and Jan Kocoń and Jauza Akbar Krito and Jebish Purbey and Jennifer Mickel and Jennifer Za and Jenny Kunz and Jihae Jeong and Jimena Tena Dávalos and Jinu Lee and João Magalhães and John Yi and Jongin Kim and Joseph Chataignon and Joseph Marvin Imperial and Jubeerathan Thevakumar and Judith Land and Junchen Jiang and Jungwhan Kim and Kairit Sirts and Kamesh R and Kamesh V and Kanda Patrick Tshinu and Kätriin Kukk and Kaustubh Ponkshe and Kavsar Huseynova and Ke He and Kelly Buchanan and Kengatharaiyer Sarveswaran and Kerem Zaman and Khalil Mrini and Kian Kyars and Krister Kruusmaa and Kusum Chouhan and Lainitha Krishnakumar and Laura Castro Sánchez and Laura Porrino Moscoso and Leshem Choshen and Levent Sencan and Lilja Øvrelid and Lisa Alazraki and Lovina Ehimen-Ugbede and Luheerathan Thevakumar and Luxshan Thavarasa and Mahnoor Malik and Mamadou K. Keita and Mansi Jangid and Marco De Santis and Marcos García and Marek Suppa and Mariam D'Ciofalo and Marii Ojastu and Maryam Sikander and Mausami Narayan and Maximos Skandalis and Mehak Mehak and Mehmet İlteriş Bozkurt and Melaku Bayu Workie and Menan Velayuthan and Michael Leventhal and Michał Marcińczuk and Mirna Potočnjak and Mohammadamin Shafiei and Mridul Sharma and Mrityunjaya Indoria and Muhammad Ravi Shulthan Habibi and Murat Kolić and Nada Galant and Naphat Permpredanun and Narada Maugin and Nicholas Kluge Corrêa and Nikola Ljubešić and Nirmal Thomas and Nisansa de Silva and Nisheeth Joshi and Nitish Ponkshe and Nizar Habash and Nneoma C. Udeze and Noel Thomas and Noémi Ligeti-Nagy and Nouhoum Coulibaly and Nsengiyumva Faustin and Odunayo Kareemat Buliaminu and Odunayo Ogundepo and Oghojafor Godswill Fejiro and Ogundipe Blessing Funmilola and Okechukwu God'spraise and Olanrewaju Samuel and Olaoye Deborah Oluwaseun and Olasoji Akindejoye and Olga Popova and Olga Snissarenko and Onyinye Anulika Chiemezie and Orkun Kinay and Osman Tursun and Owoeye Tobiloba Moses and Oyelade Oluwafemi Joshua and Oyesanmi Fiyinfoluwa and Pablo Gamallo and Pablo Rodríguez Fernández and Palak Arora and Pedro Valente and Peter Rupnik and Philip Oghenesuowho Ekiugbo and Pramit Sahoo and Prokopis Prokopidis and Pua Niau-Puhipau and Quadri Yahya and Rachele Mignone and Raghav Singhal and Ram Mohan Rao Kadiyala and Raphael Merx and Rapheal Afolayan and Ratnavel Rajalakshmi and Rishav Ghosh and Romina Oji and Ron Kekeha Solis and Rui Guerra and Rushikesh Zawar and Sa'ad Nasir Bashir and Saeed Alzaabi and Sahil Sandeep and Sai Pavan Batchu and SaiSandeep Kantareddy and Salsabila Zahirah Pranida and Sam Buchanan and Samuel Rutunda and Sander Land and Sarah Sulollari and Sardar Ali and Saroj Sapkota and Saulius Tautvaisas and Sayambhu Sen and Sayantani Banerjee and Sebastien Diarra and SenthilNathan. M and Sewoong Lee and Shaan Shah and Shankar Venkitachalam and Sharifa Djurabaeva and Sharon Ibejih and Shivanya Shomir Dutta and Siddhant Gupta and Silvia Paniagua Suárez and Sina Ahmadi and Sivasuthan Sukumar and Siyuan Song and Snegha A. and Sokratis Sofianopoulos and Sona Elza Simon and Sonja Benčina and Sophie Gvasalia and Sphurti Kirit More and Spyros Dragazis and Stephan P. Kaufhold and Suba. S and Sultan AlRashed and Surangika Ranathunga and Taiga Someya and Taja Kuzman Pungeršek and Tal Haklay and Tasi'u Jibril and Tatsuya Aoyama and Tea Abashidze and Terenz Jomar Dela Cruz and Terra Blevins and Themistoklis Nikas and Theresa Dora Idoko and Thu Mai Do and Tilek Chubakov and Tommaso Gargiani and Uma Rathore and Uni Johannesen and Uwuma Doris Ugwu and Vallerie Alexandra Putra and Vanya Bannihatti Kumar and Varsha Jeyarajalingam and Varvara Arzt and Vasudevan Nedumpozhimana and Viktoria Ondrejova and Viktoryia Horbik and Vishnu Vardhan Reddy Kummitha and Vuk Dinić and Walelign Tewabe Sewunetie and Winston Wu and Xiaojing Zhao and Yacouba Diarra and Yaniv Nikankin and Yash Mathur and Yixi Chen and Yiyuan Li and Yolanda Xavier and Yonatan Belinkov and Yusuf Ismail Abayomi and Zaid Alyafeai and Zhengyang Shan and Zhi Rui Tam and Zilu Tang and Zuzana Nadova and Baber Abbasi and Stella Biderman and David Stap and Duygu Ataman and Fabian Schmidt and Hila Gonen and Jiayi Wang and David Ifeoluwa Adelani},
year={2025},
eprint={2510.24081},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2510.24081},
}
References
See the full reference list in the paper.