Holmes: A Benchmark to Assess the Linguistic Competence of Language Models

a probing benchmark that measures what language models internally encode about linguistic phenomena

Andreas Waldis, Yotam Perlitz, Leshem Choshen, Yufang Hou, Iryna Gurevych · TACL · 2024

In one sentence

Holmes measures the linguistic competence of language models by training linear probes on the frozen last-layer representations of 208 datasets covering morphology, syntax, semantics, reasoning and discourse, isolating linguistic knowledge from instruction following.

Abstract

We introduce Holmes, a new benchmark designed to assess language models (LMs) linguistic competence - their unconscious understanding of linguistic phenomena. Specifically, we use classifier-based probing to examine LMs'internal representations regarding distinct linguistic phenomena (e.g., part-of-speech tagging). As a result, we meet recent calls to disentangle LMs'linguistic competence from other cognitive abilities, such as following instructions in prompting-based evaluations. Composing Holmes, we review over 270 probing studies and include more than 200 datasets to assess syntax, morphology, semantics, reasoning, and discourse phenomena. Analyzing over 50 LMs reveals that, aligned with known trends, their linguistic competence correlates with model size. However, surprisingly, model architecture and instruction tuning also significantly influence performance, particularly in morphology and syntax. Finally, we propose FlashHolmes, a streamlined version that reduces the computation load while maintaining high-ranking precision.

Questions this paper answers

is there a way to test how much grammar a language model actually knows without asking it questions in a prompt?
which benchmark consolidates the probing literature into a single linguistic competence suite for English language models?
where do I get a large collection of ready-made probing datasets covering morphology, syntax, semantics, reasoning and discourse?
should I use a consolidated probing benchmark instead of assembling probing datasets from individual papers myself?
Holmes is a probing benchmark that assesses the English linguistic competence of language models with 208 datasets covering 66 phenomena. Its morphology, syntax, semantics, reasoning and discourse datasets consolidate resources found in a survey of 274 probing studies.
Holds for: English only, and last-layer internal representations only. Classifier-based (linear) probing, so generation and instruction following are not measured; 18 of the 208 datasets rest on licensed resources.
A meta-study of 274 probing papers finds the literature collectively covers 289 tasks and 161 language models, yet individual studies stay narrow. Part-of-speech tagging, the most probed task, was evaluated on only 23% of the models, and the top-10 most mentioned models account for 80% of all model mentions.
Holds for: 28,063 papers from 2015 to August 2023 at ACL-family venues plus selected other venues, filtered by occurrences of 'probing'/'probe' and then manually reviewed; recent large models such as Pythia, UL2 and Llama-2 are almost absent from the surveyed work.
if a language model answers grammar questions correctly, does that mean its internal representations really encode grammar?
how closely do probing-based rankings on BLiMP agree with prompting-based rankings from HELM and the OpenLLM leaderboard?
can I skip probing and just prompt models on linguistic minimal pairs to rank their linguistic ability?
I already have leaderboard scores for my models -- do I still need probing to know their linguistic competence?
On the BLiMP datasets evaluated by both Holmes and HELM, probing-based and prompting-based model rankings barely agree, at a rank correlation of tau=0.05. Most HELM prompting results fall below the random baseline.
Holds for: 40 open decoder models and 22 BLiMP datasets covering quantifier, island effects, irregular forms and binding phenomena, using HELM's own evaluation code with its multiple-choice-joined prompting adaptation.
Holmes probing rankings correlate only moderately with the prompting-based OpenLLM leaderboard, at 54.7 to 58.0 Kendall-tau for syntax, semantics and discourse, rising to 65.3 for morphology and 77.5 for reasoning.
Holds for: Models evaluated by both Holmes and OpenLLM (Beeching et al., 2023); correlation patterns hold across MMLU, TruthfulQA and GSM8K, with TruthfulQA the weakest correlate.
do older BERT-style models capture grammar better inside than much bigger chat-style models?
how do encoder and decoder architectures compare on classifier probing for part-of-speech and agreement phenomena at matched parameter counts?
which architecture should I pick to extract reliable part-of-speech and syntactic features from hidden states?
if I need strong token-level linguistic representations, is a 70B decoder model worth it over a small encoder?
Encoder language models reach a mean winning rate of 52% on Holmes against 21% for decoder models of comparable size. Decoder models do not match encoder stability on frequent part-of-speech tokens even at 700 times the parameter count.
Holds for: 5 encoder and 6 decoder models up to 220M parameters for the mean-winning-rate comparison; token stability covers BERT, RoBERTa, GPT2, Pythia-12B and Llama-2-70B on the top-20 most common tokens of the pos, xpos and upos datasets.
does teaching a model to follow instructions change how much grammar and meaning it encodes internally?
what is the effect of instruction tuning on probing scores across morphology, syntax, semantics, reasoning and discourse?
should I probe the base checkpoint or the instruction-tuned checkpoint if I care about linguistic phenomena?
will switching to the instruction-tuned version of my model improve its internal handling of syntax and semantics?
Instruction tuning raises mean winning rate on Holmes by an average of +10% for morphology, +5% for syntax and +4% for reasoning. It lowers it by -3% for semantics and -1% for discourse, for an overall average of +2%, with per-model effects from +41% to -13%.
Holds for: 11 instruction-tuned models each compared against its own pre-trained base (Llama-2-Chat, FLAN-T5, Dolly-v2, Tülu-2, Orca-2, Vicuna-v1.5, FLAN-UL2, Mixtral-Instruct) at 7B to 70B parameters.
do bigger language models really encode more grammar than smaller ones from the same family?
how does probing performance on morphological and syntactic phenomena scale with parameter count within the Pythia and T5 families?
how big a model do I need before probing scores on syntax and morphology start to improve noticeably?
is it worth moving up to a larger checkpoint in the same family if I want better linguistic representations?
Linguistic competence on Holmes scales with parameter count within the Pythia and T5 families. The jump is pronounced beyond 0.5B parameters for Pythia and 1.0B for T5, and is concentrated in morphology and syntax.
Holds for: 8 T5 sizes and 5 Pythia sizes only; T5 (encoder-decoder) reaches a mean winning rate of 40-70% versus 20-60% for Pythia (decoder-only), so architecture and scale are not separated in this comparison.
which parts of language are models good at internally, and which parts do they handle poorly?
how do probing scores differ between formal phenomena such as morphology and syntax and functional ones such as semantics, reasoning and discourse?
which linguistic phenomena should I test if I want to find where a language model's representations are weakest?
can I assume a model that scores well on syntax probing will also handle discourse and semantics well?
Across 59 evaluated language models, linguistic competence is markedly stronger for formal phenomena (morphology and syntax) than for functional ones (semantics, reasoning and discourse), which score lower on the probing task metric.
Holds for: 59 models spanning sparse, static, encoder, decoder and encoder-decoder types, probed on last-layer representations of 208 English datasets.
The five phenomenon types in Holmes correlate at 68.4±7.5 Kendall-tau with the overall model ranking but only 54.7±13.9 with each other. Discourse is the most distinct type, at 44.4±14.7 average correlation with the others.
Holds for: Kendall-tau rank correlations over the models jointly evaluated within Holmes; semantics and reasoning correlate 73.9 and 75.6 with the overall ranking but only 58.4 with each other.
are probing scores stable enough to trust, or do they bounce around depending on how you run them?
how much do classifier probing results vary across random seeds compared with the variance from prompt paraphrasing?
how many seeds do I need to run before probing numbers are reliable enough to compare models?
should I trust a probing-based ranking of my models, or will it change if I rerun it?
Probing results on Holmes vary little across seeds, with an average standard deviation of 0.02 over 5 random seeds, against the 0.07 deviation reported for prompt paraphrasing in prompting-based evaluation. Average compression is 1.9 and average selectivity 0.31.
Holds for: Averages over 208 probing datasets and the evaluated models, with fixed probe hyperparameters (20 epochs, batch size 64, learning rate 0.0005); selectivity computed only for base-sized models of 10M-200M parameters.
how much computing time does it take to run a full battery of internal linguistic tests on a large language model?
what is the compute cost of encoding 208 probing datasets for a 70B model, and how much does a subsampled variant save?
how do I evaluate a newly released model on a large probing suite without spending GPU days on encoding?
on a small compute budget, can I use a cheap subsampled probing run and still get the same model ordering?
FlashHolmes reproduces Holmes model rankings at a rank resolution of about 1.5 while requiring roughly 3% of the computation. It trains probes on 1/32 of the training instances and drops the 18 licensed datasets, versus about 6 GPU days of encoding for a 70B model on full Holmes.
Holds for: Rank resolution is the 95% CI of rank difference against full Holmes, where 1/2 of the training data gives about 0.9 and 1/512 about 2.6; English datasets, last-layer representations.
how much do published studies of what language models know about language actually overlap in the tasks and models they test?
how many probing tasks and language models does the probing literature jointly cover, and what share uses classifier-based versus mask-based probing?
how do I find out whether the probing task I care about has already been evaluated on the models I use?
can I rely on published probing papers to tell me about my model, or do they each cover too few models?
A meta-study of 274 probing papers finds the literature collectively covers 289 tasks and 161 language models, yet individual studies stay narrow. Part-of-speech tagging, the most probed task, was evaluated on only 23% of the models, and the top-10 most mentioned models account for 80% of all model mentions.
Holds for: 28,063 papers from 2015 to August 2023 at ACL-family venues plus selected other venues, filtered by occurrences of 'probing'/'probe' and then manually reviewed; recent large models such as Pythia, UL2 and Llama-2 are almost absent from the surveyed work.
Among 274 surveyed probing studies, 74% use classifier-based probing and 20% use mask-based probing, while roughly 3% rely on attention patterns or other approaches.
Holds for: Categorisation of studies published 2015 to August 2023 at major NLP venues; each study assigned a single dominant probing method.
if a test set ended up in a model's training data, does that ruin an evaluation of its internal linguistic knowledge?
why is probing internal representations argued to remain valid under benchmark contamination in pretraining corpora?
how do I evaluate linguistic competence when I cannot rule out that the evaluation datasets were seen during pretraining?
should I worry about data leakage if I evaluate my model with probing classifiers rather than prompts?
Holmes argues that probing-based evaluation retains validity under dataset contamination, because instruction tuning aligns a model's textual responses rather than explicitly aligning the internal representations that the probes read.
Holds for: An argument rather than a measurement; training data is unknown for several evaluated models, including Llama-2, Mixtral and Wizard.
what exactly does a large linguistic probing benchmark measure about a language model, and how much does it cover?
which linguistic phenomena and dataset counts make up the Holmes probing suite, and what does it find about formal versus functional competence?
what will I learn about my model's linguistic abilities if I run it through a broad probing benchmark?
is a broad linguistic probing benchmark measuring the phenomena I care about before I commit to running it?
Holmes is a probing benchmark that assesses the English linguistic competence of language models with 208 datasets covering 66 phenomena. Its morphology, syntax, semantics, reasoning and discourse datasets consolidate resources found in a survey of 274 probing studies.
Holds for: English only, and last-layer internal representations only. Classifier-based (linear) probing, so generation and instruction following are not measured; 18 of the 208 datasets rest on licensed resources.
Across 59 evaluated language models, linguistic competence is markedly stronger for formal phenomena (morphology and syntax) than for functional ones (semantics, reasoning and discourse), which score lower on the probing task metric.
Holds for: 59 models spanning sparse, static, encoder, decoder and encoder-decoder types, probed on last-layer representations of 208 English datasets.

Claims and scope

Common misreadings

Terminology in this paper

linguistic competence
Following Chomsky (1965), a language model's unconscious internal understanding of linguistic phenomena, assessed via its internal representations rather than via its textual responses.
linguistic performance
A language model's use of language in textual responses to instructions, which is what prompting-based benchmarks measure and which conflates linguistic knowledge with abilities such as instruction following.
formal vs. functional phenomena
Formal phenomena are morphology and syntax (grammatical rules and statistical patterns); functional phenomena are semantics, reasoning and discourse (practical abilities like interpreting sentiment or detecting speculation).
selectivity
The macro-F1 of a probe trained on the true labels minus the macro-F1 of the same probe trained on randomly assigned control-task labels; higher values mean the probe found phenomenon-relevant structure rather than memorising.
compression
The ratio of a uniform encoding of instances and labels to their minimum description length under the probe; higher values mean the linguistic phenomenon is more cleanly encoded in the representation.
discriminability
The Kendall-tau alignment between the model ranking induced by a single probing dataset and the overall benchmark ranking; low values mean no single dataset dominates the aggregate ranking.
rank resolution
The 95% confidence interval of the difference in a model's rank between a subsampled benchmark and the full benchmark; a resolution of 1 means a model keeps its rank or swaps with a neighbour.
FlashHolmes
The streamlined variant of the Holmes benchmark that trains probes on 1/32 of the training instances and excludes 18 licensed datasets.

How to cite

@article{DBLP:journals/corr/abs-2404-18923,author       = {Andreas Waldis and
                  Yotam Perlitz and
                  Leshem Choshen and
                  Yufang Hou and
                  Iryna Gurevych},
  title        = {Holmes: {A} Benchmark to Assess the Linguistic Competence of Language Models},
  journal      = {Trans. Assoc. Comput. Linguistics},
  volume       = {12},
  year         = {2024},
  url          = {https://doi.org/10.1162/tacl\_a\_00718},
  doi          = {10.1162/TACL\_A\_00718},
  eprinttype    = {arXiv},
  eprint       = {2404.18923},
  timestamp    = {Tue, 14 Oct 2025 01:00:00 +0200},
  biburl       = {https://dblp.org/rec/journals/tacl/WaldisWPCCHG24.bib},
  bibsource    = {dblp computer science bibliography, https://dblp.org},
  pages        = {1616--1647}
}

References

See the full reference list in the paper.