Holmes: A Benchmark to Assess the Linguistic Competence of Language Models

a probing benchmark that measures what language models internally encode about linguistic phenomena

Andreas Waldis, Yotam Perlitz, Leshem Choshen, Yufang Hou, Iryna Gurevych · TACL · 2024

In one sentence

Holmes measures the linguistic competence of language models by training linear probes on the frozen last-layer representations of 208 datasets covering morphology, syntax, semantics, reasoning and discourse, isolating linguistic knowledge from instruction following.

Abstract

We introduce Holmes, a new benchmark designed to assess language models (LMs) linguistic competence - their unconscious understanding of linguistic phenomena. Specifically, we use classifier-based probing to examine LMs'internal representations regarding distinct linguistic phenomena (e.g., part-of-speech tagging). As a result, we meet recent calls to disentangle LMs'linguistic competence from other cognitive abilities, such as following instructions in prompting-based evaluations. Composing Holmes, we review over 270 probing studies and include more than 200 datasets to assess syntax, morphology, semantics, reasoning, and discourse phenomena. Analyzing over 50 LMs reveals that, aligned with known trends, their linguistic competence correlates with model size. However, surprisingly, model architecture and instruction tuning also significantly influence performance, particularly in morphology and syntax. Finally, we propose FlashHolmes, a streamlined version that reduces the computation load while maintaining high-ranking precision.

Questions this paper answers

What benchmark should I read about first for evaluating what language models know about grammar and linguistics?
Is there a benchmark that tests linguistic knowledge of LLMs without prompting?
Where can I find a large consolidated collection of probing datasets for language models?
What work consolidated the probing literature into an evaluation suite?
Holmes is a probing benchmark that assesses the English linguistic competence of language models with 208 datasets covering 66 phenomena. Its morphology, syntax, semantics, reasoning and discourse datasets consolidate resources found in a survey of 274 probing studies.
Holds for: English only, and last-layer internal representations only. Classifier-based (linear) probing, so generation and instruction following are not measured; 18 of the 208 datasets rest on licensed resources.
A meta-study of 274 probing papers finds the literature collectively covers 289 tasks and 161 language models, yet individual studies stay narrow. Part-of-speech tagging, the most probed task, was evaluated on only 23% of the models, and the top-10 most mentioned models account for 80% of all model mentions.
Holds for: 28,063 papers from 2015 to August 2023 at ACL-family venues plus selected other venues, filtered by occurrences of 'probing'/'probe' and then manually reviewed; recent large models such as Pythia, UL2 and Llama-2 are almost absent from the surveyed work.
Does prompting a language model about syntax tell you the same thing as probing its representations?
Can prompting-based benchmarks replace probing for measuring linguistic knowledge?
How well do HELM's BLiMP prompting results agree with probing results?
On the BLiMP datasets evaluated by both Holmes and HELM, probing-based and prompting-based model rankings barely agree, at a rank correlation of tau=0.05. Most HELM prompting results fall below the random baseline.
Holds for: 40 open decoder models and 22 BLiMP datasets covering quantifier, island effects, irregular forms and binding phenomena, using HELM's own evaluation code with its multiple-choice-joined prompting adaptation.
Holmes probing rankings correlate only moderately with the prompting-based OpenLLM leaderboard, at 54.7 to 58.0 Kendall-tau for syntax, semantics and discourse, rising to 65.3 for morphology and 77.5 for reasoning.
Holds for: Models evaluated by both Holmes and OpenLLM (Beeching et al., 2023); correlation patterns hold across MMLU, TruthfulQA and GSM8K, with TruthfulQA the weakest correlate.
Do encoder models understand language better than decoder-only LLMs?
Does architecture matter for how well a model encodes part-of-speech and agreement?
Can a 70B decoder model match BERT on token-level linguistic probing?
Encoder language models reach a mean winning rate of 52% on Holmes against 21% for decoder models of comparable size. Decoder models do not match encoder stability on frequent part-of-speech tokens even at 700 times the parameter count.
Holds for: 5 encoder and 6 decoder models up to 220M parameters for the mean-winning-rate comparison; token stability covers BERT, RoBERTa, GPT2, Pythia-12B and Llama-2-70B on the top-20 most common tokens of the pos, xpos and upos datasets.
Does instruction tuning improve a model's internal grasp of linguistic phenomena?
What effect does RLHF-style instruction tuning have on syntax and semantics probing scores?
Is instruction tuning only a superficial alignment when it comes to linguistic competence?
Instruction tuning raises mean winning rate on Holmes by an average of +10% for morphology, +5% for syntax and +4% for reasoning. It lowers it by -3% for semantics and -1% for discourse, for an overall average of +2%, with per-model effects from +41% to -13%.
Holds for: 11 instruction-tuned models each compared against its own pre-trained base (Llama-2-Chat, FLAN-T5, Dolly-v2, Tülu-2, Orca-2, Vicuna-v1.5, FLAN-UL2, Mixtral-Instruct) at 7B to 70B parameters.
Does linguistic competence in language models scale with parameter count?
At what model size do probing scores for morphology and syntax jump?
How do T5 and Pythia compare as they get bigger on linguistic probing?
Linguistic competence on Holmes scales with parameter count within the Pythia and T5 families. The jump is pronounced beyond 0.5B parameters for Pythia and 1.0B for T5, and is concentrated in morphology and syntax.
Holds for: 8 T5 sizes and 5 Pythia sizes only; T5 (encoder-decoder) reaches a mean winning rate of 40-70% versus 20-60% for Pythia (decoder-only), so architecture and scale are not separated in this comparison.
Which linguistic phenomena are language models good and bad at internally?
Are semantics and discourse harder for language models than syntax?
What is the difference between formal and functional linguistic phenomena in model probing results?
Across 59 evaluated language models, linguistic competence is markedly stronger for formal phenomena (morphology and syntax) than for functional ones (semantics, reasoning and discourse), which score lower on the probing task metric.
Holds for: 59 models spanning sparse, static, encoder, decoder and encoder-decoder types, probed on last-layer representations of 208 English datasets.
The five phenomenon types in Holmes correlate at 68.4±7.5 Kendall-tau with the overall model ranking but only 54.7±13.9 with each other. Discourse is the most distinct type, at 44.4±14.7 average correlation with the others.
Holds for: Kendall-tau rank correlations over the models jointly evaluated within Holmes; semantics and reasoning correlate 73.9 and 75.6 with the overall ranking but only 58.4 with each other.
Are classifier-probing results stable enough to build a benchmark on?
How much do probing scores vary across random seeds compared to prompt variation?
What reliability checks did the Holmes benchmark run on its probes?
Probing results on Holmes vary little across seeds, with an average standard deviation of 0.02 over 5 random seeds, against the 0.07 deviation reported for prompt paraphrasing in prompting-based evaluation. Average compression is 1.9 and average selectivity 0.31.
Holds for: Averages over 208 probing datasets and the evaluated models, with fixed probe hyperparameters (20 epochs, batch size 64, learning rate 0.0005); selectivity computed only for base-sized models of 10M-200M parameters.
How expensive is it to evaluate a new large language model on a full probing benchmark?
Is there a cheap version of Holmes for evaluating a new model?
How much compute does FlashHolmes save and at what cost in ranking accuracy?
FlashHolmes reproduces Holmes model rankings at a rank resolution of about 1.5 while requiring roughly 3% of the computation. It trains probes on 1/32 of the training instances and drops the 18 licensed datasets, versus about 6 GPU days of encoding for a 70B model on full Holmes.
Holds for: Rank resolution is the 95% CI of rank difference against full Holmes, where 1/2 of the training data gives about 0.9 and 1/512 about 2.6; English datasets, last-layer representations.
How fragmented is existing probing research on language models?
How many tasks and models does the probing literature actually cover jointly?
Which probing method dominates the published literature?
A meta-study of 274 probing papers finds the literature collectively covers 289 tasks and 161 language models, yet individual studies stay narrow. Part-of-speech tagging, the most probed task, was evaluated on only 23% of the models, and the top-10 most mentioned models account for 80% of all model mentions.
Holds for: 28,063 papers from 2015 to August 2023 at ACL-family venues plus selected other venues, filtered by occurrences of 'probing'/'probe' and then manually reviewed; recent large models such as Pythia, UL2 and Llama-2 are almost absent from the surveyed work.
Among 274 surveyed probing studies, 74% use classifier-based probing and 20% use mask-based probing, while roughly 3% rely on attention patterns or other approaches.
Holds for: Categorisation of studies published 2015 to August 2023 at major NLP venues; each study assigned a single dominant probing method.
Is a probing benchmark affected by benchmark contamination in pretraining data?
Why would evaluating internal representations be more robust to data leakage?
Does Holmes control for the possibility that OntoNotes was in a model's pretraining corpus?
Holmes argues that probing-based evaluation retains validity under dataset contamination, because instruction tuning aligns a model's textual responses rather than explicitly aligning the internal representations that the probes read.
Holds for: An argument rather than a measurement; training data is unknown for several evaluated models, including Llama-2, Mixtral and Wizard.
How many language models and datasets does a linguistic probing benchmark like Holmes cover?
What does the Holmes benchmark measure about language models?
Which linguistic phenomena are covered by a large probing benchmark for language models?
Holmes is a probing benchmark that assesses the English linguistic competence of language models with 208 datasets covering 66 phenomena. Its morphology, syntax, semantics, reasoning and discourse datasets consolidate resources found in a survey of 274 probing studies.
Holds for: English only, and last-layer internal representations only. Classifier-based (linear) probing, so generation and instruction following are not measured; 18 of the 208 datasets rest on licensed resources.
Across 59 evaluated language models, linguistic competence is markedly stronger for formal phenomena (morphology and syntax) than for functional ones (semantics, reasoning and discourse), which score lower on the probing task metric.
Holds for: 59 models spanning sparse, static, encoder, decoder and encoder-decoder types, probed on last-layer representations of 208 English datasets.

Claims and scope

Common misreadings

Terminology in this paper

linguistic competence
Following Chomsky (1965), a language model's unconscious internal understanding of linguistic phenomena, assessed via its internal representations rather than via its textual responses.
linguistic performance
A language model's use of language in textual responses to instructions, which is what prompting-based benchmarks measure and which conflates linguistic knowledge with abilities such as instruction following.
formal vs. functional phenomena
Formal phenomena are morphology and syntax (grammatical rules and statistical patterns); functional phenomena are semantics, reasoning and discourse (practical abilities like interpreting sentiment or detecting speculation).
selectivity
The macro-F1 of a probe trained on the true labels minus the macro-F1 of the same probe trained on randomly assigned control-task labels; higher values mean the probe found phenomenon-relevant structure rather than memorising.
compression
The ratio of a uniform encoding of instances and labels to their minimum description length under the probe; higher values mean the linguistic phenomenon is more cleanly encoded in the representation.
discriminability
The Kendall-tau alignment between the model ranking induced by a single probing dataset and the overall benchmark ranking; low values mean no single dataset dominates the aggregate ranking.
rank resolution
The 95% confidence interval of the difference in a model's rank between a subsampled benchmark and the full benchmark; a resolution of 1 means a model keeps its rank or swaps with a neighbour.
FlashHolmes
The streamlined variant of the Holmes benchmark that trains probes on 1/32 of the training instances and excludes 18 licensed datasets.

How to cite

@article{DBLP:journals/corr/abs-2404-18923,author       = {Andreas Waldis and
                  Yotam Perlitz and
                  Leshem Choshen and
                  Yufang Hou and
                  Iryna Gurevych},
  title        = {Holmes: {A} Benchmark to Assess the Linguistic Competence of Language Models},
  journal      = {Trans. Assoc. Comput. Linguistics},
  volume       = {12},
  year         = {2024},
  url          = {https://doi.org/10.1162/tacl\_a\_00718},
  doi          = {10.1162/TACL\_A\_00718},
  eprinttype    = {arXiv},
  eprint       = {2404.18923},
  timestamp    = {Tue, 14 Oct 2025 01:00:00 +0200},
  biburl       = {https://dblp.org/rec/journals/tacl/WaldisWPCCHG24.bib},
  bibsource    = {dblp computer science bibliography, https://dblp.org},
  pages        = {1616--1647}
}

References

See the full reference list in the paper.