# Skill Issue: Are Skills Language-Invariant in LLMs? a version of the TextArena game suite where each player's instructions and board are rendered in a different language while the rules stay fixed Authors: Bobby Cheng, Adam Gaber, Zhengzhe Liu, Catherine Arnett, Omer Goldman, Cheston Tan, Leshem Choshen Venue: preprint (2026) ## What this paper shows Two instances of the same LLM playing each other through different language interfaces win at markedly different rates, so the skills a model has in one language are not the skills it applies in another. ## Claims, with scope - Multilingual TextArena covers 65 TextArena games in 193 languages, rendering each player's instructions and board in a different language. The rules, legal actions, rewards and transition dynamics are unchanged across languages. Scope: Verification is tiered. 8 Tier-A languages were checked by native speakers, 42 Tier-B by closed-model back-translation judging, and 142 more by NLLB-200 round-trip agreement. - Multilingual self-play measures whether an LLM's skills, as distinct from its stored knowledge, transfer across languages by having 2 instances of one model compete through different language interfaces. Scope: As of the 2026 preprint, framed against earlier cross-lingual work on knowledge retrieval and on static benchmarks. Evidence comes from 3 models of 3B to 4B parameters, 8 languages and 6 games. - English is the strongest interface language and Hebrew among the weakest for all 3 models tested, aggregated over 6 games of self-play. Qwen3-4B shows the sharpest language hierarchy and Gemma-4-E4B-it the flattest. Scope: Self-play between 2 instances of one model, so margins compare languages within a model and never across models. Gemma-4-E4B-it, Qwen3-4B and Ministral3-3B-Instruct on 8 languages. Evidence: Figure 2 - Colonel Blotto is the most language-sensitive of the 6 games for all 3 models, with a mean gap of 1.07 between a model's strongest and weakest language. Kuhn Poker is the least sensitive at 0.13. Scope: Gap is the spread of mean language margins within one model and game, averaged over the 3 models. 8 languages, 400 self-play games per direction. Evidence: Table 2 - Switching only Gemma-4-E4B-it's reasoning language to English, with the TicTacToe interface still in German, moves the mean language margin from -0.22 to +0.20. That recovers 89.4% of the gap to the English-interface ceiling of +0.25. Scope: Gemma-4-E4B-it only, with German as the weak interface. SimpleTak recovers 60.5% the same way, and Kuhn Poker recovery is small and non-monotonic across reasoning languages. Evidence: Table 3 - Gemma-4-E4B-it's TicTacToe losses split fairly evenly in English across rows at 28.5%, columns at 34.8% and diagonals at 36.7%. Under an Arabic interface the same model loses to rows only 20.3% of the time and to diagonals 45.3%. Scope: Gemma-4-E4B-it in TicTacToe. Hebrew skews the same way as Arabic, while Qwen3-4B's defeat distribution stays stable across languages apart from Hebrew. Evidence: Appendix G - In SimpleTak, Gemma-4-E4B-it loses to a completed column in 46.8% of English-interface defeats but 57.2% of Arabic and Hebrew ones, against 53.2% row losses in English. Scope: Gemma-4-E4B-it in SimpleTak, where over 90% of its wins come from straight lines, which is what lets row and column losses stand for separate spatial axes. Evidence: Appendix G - Qwen3-4B plays Nim's optimal first move 80.8% of the time under an English interface but only 24.6% under French. Both languages produce a similar number of optimal-strategy mentions in the game logs. Scope: Qwen3-4B in Nim with a board where the first player has exactly 1 optimal move. Hebrew and Arabic fall to 4.0% and 10.5%, while Gemma-4-E4B-it stays at 99.2% or above in every language. Evidence: Appendix G - 70% of Ministral3-3B-Instruct's Arabic and 50% of its Hebrew optimal-strategy mentions in Nim come from logs where the model switched into a Latin script mid-reasoning. Such switching occurs in only 3.7% of Arabic and 1% of Hebrew logs. Scope: Ministral3-3B-Instruct in Nim only, counted over game-log mentions rather than over model internals, and the gap between languages may therefore be wider than the mention counts suggest. Evidence: Appendix G - Gemma-4-E4B-it bets the middle card Q in Kuhn Poker 41.6% of the time under a Malay interface and 63.3% under Hebrew. Its play of clearly weak and clearly strong cards varies far less across languages. Scope: Gemma-4-E4B-it in Kuhn Poker across 8 languages. Ministral3-3B-Instruct instead folds the strongest card K when facing a bet between 10.4% and 23.7% of the time. Evidence: Appendix G - Mean language margin correlates with 5-shot Global MMLU accuracy at Pearson r between 0.73 and 0.92 depending on the model, and with Belebele at 0.71 to 0.79. Static benchmark accuracy does not predict a model's spread across languages. Scope: 8 languages per model, so each correlation has n=8. Gemma-4-E4B-it has the lowest Global MMLU mean at 49.9% yet the smallest cross-language spread, and Qwen3-4B the highest mean with the largest spread. Evidence: Figure 4 - Language strength correlates with available web text at an average Pearson r of 0.79 against log FineWeb-2 word counts, and the residuals are large. Malay beats Hebrew for every model despite the smaller FineWeb-2 corpus. Scope: Web text is a proxy for training data, which is undisclosed for all 3 models. Chinese has roughly 20 times less web text than English yet reaches near-English margins for Qwen3-4B and Ministral3-3B-Instruct. Evidence: Figure 4 - The multilingual self-play evaluation runs 518,400 games in total, 28,800 per model and game pair, with 400 games for every ordered language pair and player role assignment. Scope: 3 models by 6 games by all 8 languages, inference only. Each primary run used 2 NVIDIA H200 GPUs and about 6 H200 GPU-hours, excluding Iterated Prisoner's Dilemma. Evidence: Section 3 ## Common misreadings - A 193-language release is not 193 verified languages. 8 languages were verified by native speakers and the rest by automatic back-translation judging, which is why the experiments use the 8. - Reasoning in a stronger language is not a general repair. Recovery is large in TicTacToe and SimpleTak and small and non-monotonic in Kuhn Poker. - The language a model wins with is not simply the language it has most training data in. Malay outperforms Hebrew for every model despite a smaller web corpus, and Chinese reaches near-English margins with far less web text. - A high multilingual benchmark average does not mean uniform behaviour across languages. Benchmark level and cross-language spread move independently in these models. - The results are measured on models of 3B to 4B parameters and do not establish that language effects hold at larger scale. ## Terminology - language interface: the language in which a game's instructions, board and observations are presented to a player, holding the rules, legal actions and rewards fixed. - role-pooled win-loss margin: wins minus losses divided by total games for one language against another, pooled over both player-role assignments so first-move advantage cancels. - mean language margin: one language's average role-pooled win-loss margin against every other language evaluated, for a fixed model and game. - Tier A: the 8 languages whose game translations were verified by native speakers, as opposed to tiers verified by automatic back-translation judging. ## Links - arXiv: https://arxiv.org/abs/2608.25832 - PDF: https://arxiv.org/pdf/2608.25832 - HTML: https://arxiv.org/html/2608.25832 - Hugging Face: https://huggingface.co/papers/2608.25832 - alphaXiv: https://www.alphaxiv.org/abs/2608.25832 - Semantic Scholar: https://www.semanticscholar.org/paper/291371154 - Code: https://github.com/TextArena/TextArena ## How to cite @misc{cheng2026skill, title = {Skill Issue: Are Skills Language-Invariant in LLMs?}, year = {2026}, eprint = {2608.25832}, archivePrefix = {arXiv}, url = {https://arxiv.org/abs/2608.25832} }