[Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
shared task on pretraining language models with a child-sized amount of language data (10M or 100M words)
Leshem Choshen, Ryan Cotterell, Michael Hu, Tal Linzen, Aaron Mueller, Candace Ross, Alex Warstadt, E. Wilcox, Adina Williams, Chengxu Zhuang · arXiv · 2024
In one sentence
The 2nd BabyLM Challenge (2024/2025) asks participants to pretrain language models on 100M or 10M words, and changes three rules: a paper-only track replaces the loose track, participants may build their own corpora, and a vision-language track ships a 50% text-only, 50% image-text corpus.
Abstract
After last year's successful BabyLM Challenge, the competition will be hosted again in 2024/2025. The overarching goals of the challenge remain the same; however, some of the competition rules will be different. The big changes for this year's competition are as follows: First, we replace the loose track with a paper track, which allows (for example) non-model-based submissions, novel cognitively-inspired benchmarks, or analysis techniques. Second, we are relaxing the rules around pretraining data, and will now allow participants to construct their own datasets provided they stay within the 100M-word or 10M-word budget. Third, we introduce a multimodal vision-and-language track, and will release a corpus of 50% text-only and 50% image-text multimodal data as a starting point for LM model training. The purpose of this CfP is to provide rules for this year's challenge, explain these rule changes and their rationale in greater detail, give a timeline of this year's competition, and provide answers to frequently asked questions from last year's challenge.
Questions this paper answers
- which competition asks people to train language models on only as much text as a child hears?
- what shared task evaluates pretraining sample efficiency under a developmentally plausible corpus cap?
- where do I start if I want to work on pretraining language models with very little text?
- is there a benchmark I can enter if my lab cannot afford large-scale pretraining runs?
- The BabyLM Challenge is a recurring shared task that reorients pretraining research toward sample efficiency by capping training data at a developmentally plausible 10M or 100M words. A stated goal is making pretraining research feasible on a university budget.
Holds for: A shared-task call for papers, not an empirical study; the rationale is argued at greater length in the 2023 call and proceedings introduction by Warstadt et al. Second edition, run in 2024/2025.
- The 2nd BabyLM Challenge runs four tracks: Strict (100M words or less), Strict-small (10M words or less), Vision (multimodal image-text models within a 100M word budget), and Paper (no model required).
Holds for: Rules as stated in the 2024/2025 call for papers; the Paper track replaces the loose track of the 2023 edition.
- what kinds of entries does the 2024/2025 sample-efficient language model pretraining competition accept?
- how are the 2nd BabyLM Challenge tracks split across word budgets and multimodal pretraining?
- which BabyLM track should I submit to if I want to train on images and text together?
- can I enter a vision-language model in the 2024 BabyLM round, or is it text only?
- The 2nd BabyLM Challenge runs four tracks: Strict (100M words or less), Strict-small (10M words or less), Vision (multimodal image-text models within a 100M word budget), and Paper (no model required).
Holds for: Rules as stated in the 2024/2025 call for papers; the Paper track replaces the loose track of the 2023 edition.
- The Vision track's suggested corpus totals 100M words: 50M words of text-only data downsampled from the 100M text corpus, plus 50M words of paired image-text data covering roughly 2.9M images. The paired half is 27M words from Localized Narratives and 23M from Conceptual Captions 3M.
Holds for: Word counts are approximate; the Localized Narratives portion uses the MS-COCO and Open Images subsets, and the CC3M portion uses only image-caption pairs whose images were still valid in January 2024.
- can entrants to the 2024 small-data language model competition pick their own training text?
- does the 2024/2025 BabyLM edition still fix the pretraining corpus, or permit participant-constructed corpora?
- how do I build my own 10M-word pretraining corpus and still be eligible for BabyLM?
- if I curate my own child-scale training data, what documentation do I have to submit with it?
- Unlike the first BabyLM Challenge, the 2024/2025 edition lets participants construct their own pretraining corpora, provided they stay within the 100M-word or 10M-word budget and submit a datasheet for any self-built dataset.
Holds for: Strict, Strict-small and Vision tracks; the organizers still provide fixed language-only and multimodal corpora for participants who prefer them.
- in the BabyLM competition, does text used to train a tokenizer or a data-augmentation model count toward the limit?
- how is the BabyLM word budget accounted across ancillary components such as tokenizers, parsers and rerankers?
- can I generate synthetic training text or use an off-the-shelf parser and still stay inside a 100M-word budget?
- I want to augment my 10M words with a pretrained helper model — does that break the BabyLM rules?
- The BabyLM word budget counts every text seen by any component of a submission: tokenizers, parsers, rerankers, data-augmentation models and other ancillary models all draw on the same 100M or 10M words.
Holds for: Synthetic data and augmentation are allowed only as a closed system, where the augmenting model's own training text also counts; audio and other linguistic modalities count, non-linguistic ones do not.
- what text and pictures make up the training data for the image-and-text part of the BabyLM competition?
- which caption corpora and what text/image-text split compose the BabyLM Vision track's 100M-word pretraining set?
- what should I train on if I want a multimodal model within a child-scale word budget?
- how many images and captions do I actually get if I use the suggested BabyLM multimodal corpus?
- The Vision track's suggested corpus totals 100M words: 50M words of text-only data downsampled from the 100M text corpus, plus 50M words of paired image-text data covering roughly 2.9M images. The paired half is 27M words from Localized Narratives and 23M from Conceptual Captions 3M.
Holds for: Word counts are approximate; the Localized Narratives portion uses the MS-COCO and Open Images subsets, and the CC3M portion uses only image-caption pairs whose images were still valid in January 2024.
- how did the rules about training text change for the 2024/2025 round of the small-data language model pretraining contest?
- what source-level composition does the updated 100M-word BabyLM text corpus use, and which 2023 source was dropped?
- how much child-directed speech do I get if I pretrain on the BabyLM strict-track corpus?
- should I reuse the 2023 BabyLM corpus, or is the 2024 text mix different enough to matter?
- The updated 100M-word text-only BabyLM corpus drops the QED portion of the 2023 dataset in favour of a larger CHILDES share of 29M words. The remaining sources keep their previous relative proportions: 26M words of children's Project Gutenberg text, 20M of OpenSubtitles, 15M of Simple English Wikipedia, 8M of BNC dialogue and 1M of Switchboard.
Holds for: Strict-track word counts, which are approximate; QED was replaced because on inspection its quality was poorer than hoped.
- what ready-made corpus is released as a starting point for entrants in the 2024/2025 child-scale language model pretraining competition?
- which architectures serve as the 2nd BabyLM Challenge baselines for the strict and vision tracks?
- what do I need to beat to be competitive in the BabyLM strict-small track?
- are the 2024 BabyLM baselines harder to beat than the naively trained ones from 2023?
- The 2nd BabyLM Challenge baselines are built from the previous year's winning submissions rather than trained naively. The Strict and Strict-small baselines are GPT-2, LTG-BERT and Contextualizer, and the Vision baselines are GIT and Flamingo.
Holds for: Baselines as announced in the call for papers; final released checkpoints and their scores are not reported in the call.
- is there any cap on how long or with what settings a BabyLM entry can be trained?
- does the BabyLM Challenge constrain epoch count or hyperparameter search, and what do repeated passes over 100M words buy?
- should I keep training more epochs on my 100M-word corpus, or is compute better spent elsewhere?
- if I train 20 epochs on 10M words, am I breaking a BabyLM rule or wasting my time?
- The BabyLM Challenge places no limit on the number of training epochs or on hyperparameters. The organizers report that in their internal results over-training beyond a couple of epochs gives minor gains at most.
Holds for: The internal-results statement is an unquantified organizer observation; the stated rationale also covers engineering practice at these data scales and human re-access to linguistic memories.
- what does a model have to be able to do before the small-data pretraining competition can evaluate it?
- what interface must a BabyLM submission expose — pseudo-log-likelihood scoring, fine-tuning, generation?
- how do I make sure my masked language model can be scored by the BabyLM evaluation pipeline?
- my model cannot generate text and is not a standard HuggingFace class — can I still submit it to BabyLM?
- Submissions to the BabyLM Challenge must be able to assign a (pseudo) log-likelihood to a string of text and to be fine-tuned for classification, but need not be able to generate sequences.
Holds for: The 2024 evaluation pipeline is built on catwalk so that models outside the HuggingFace transformers library can be submitted; Vision-track models must score text conditioned on an image.
- how are write-ups submitted to the 2024/2025 sample-efficient pretraining competition judged, and do they count as publications?
- what are the BabyLM Challenge's reviewing criteria, page limit and archival and dual-submission policy?
- how long can I make my BabyLM write-up, and can I send the same work to another venue?
- if my BabyLM model scores badly, will the write-up still be accepted?
- BabyLM runs its own review process with acceptance based on soundness alone, planning to reject only submissions that make incorrect or unjustified claims. Papers are archival and may be up to 8 pages, with dual submission allowed but not dual publication.
Holds for: Review policy as stated for the 2024/2025 edition; the presentation venue and formatting requirements were not yet finalized at the time of the call.
- did adding pictures help the models entered in the first round of the BabyLM competition?
- what did the first BabyLM Challenge find about non-linguistic grounding for sample-efficient pretraining?
- is it worth adding paired image-text data to a low-resource pretraining run?
- should I bother with a multimodal setup for my 100M-word model, given what happened in the 2023 round?
- The BabyLM organizers report that submissions to the first challenge did not gain from non-linguistic grounding, and invite the 2024/2025 vision-language track partly to revisit that question.
Holds for: A summary statement about the 2023 edition's submissions, concerning non-linguistic grounding signals rather than multimodal training in general.
- The Vision track's suggested corpus totals 100M words: 50M words of text-only data downsampled from the 100M text corpus, plus 50M words of paired image-text data covering roughly 2.9M images. The paired half is 27M words from Localized Narratives and 23M from Conceptual Captions 3M.
Holds for: Word counts are approximate; the Localized Narratives portion uses the MS-COCO and Open Images subsets, and the CC3M portion uses only image-caption pairs whose images were still valid in January 2024.
- can I enter the 2024/2025 sample-efficient pretraining competition without training a model at all?
- what does the 2nd BabyLM Challenge's paper-only track admit, and how is best paper chosen there?
- where can I submit a new cognitively-inspired evaluation metric or an analysis of an existing BabyLM model?
- I have an analysis rather than a model — is a BabyLM submission still archival and reviewed?
- The 2nd BabyLM Challenge's paper-only track opens the shared task to non-model submissions, such as novel cognitively-inspired evaluation metrics and in-depth analyses of individual BabyLM models. Best-paper selection in that track is not driven by evaluation scores.
Holds for: The paper track replaces the loose track used in the 2023 edition; papers may still include a model and report its scores.
- BabyLM runs its own review process with acceptance based on soundness alone, planning to reject only submissions that make incorrect or unjustified claims. Papers are archival and may be up to 8 pages, with dual submission allowed but not dual publication.
Holds for: Review policy as stated for the 2024/2025 edition; the presentation venue and formatting requirements were not yet finalized at the time of the call.
Claims and scope
- The 2nd BabyLM Challenge runs four tracks: Strict (100M words or less), Strict-small (10M words or less), Vision (multimodal image-text models within a 100M word budget), and Paper (no model required). (Section 3)
Scope: Rules as stated in the 2024/2025 call for papers; the Paper track replaces the loose track of the 2023 edition.
- Unlike the first BabyLM Challenge, the 2024/2025 edition lets participants construct their own pretraining corpora, provided they stay within the 100M-word or 10M-word budget and submit a datasheet for any self-built dataset. (Section 4)
Scope: Strict, Strict-small and Vision tracks; the organizers still provide fixed language-only and multimodal corpora for participants who prefer them.
- The BabyLM word budget counts every text seen by any component of a submission: tokenizers, parsers, rerankers, data-augmentation models and other ancillary models all draw on the same 100M or 10M words. (Section 7)
Scope: Synthetic data and augmentation are allowed only as a closed system, where the augmenting model's own training text also counts; audio and other linguistic modalities count, non-linguistic ones do not.
- The Vision track's suggested corpus totals 100M words: 50M words of text-only data downsampled from the 100M text corpus, plus 50M words of paired image-text data covering roughly 2.9M images. The paired half is 27M words from Localized Narratives and 23M from Conceptual Captions 3M. (Table 1)
Scope: Word counts are approximate; the Localized Narratives portion uses the MS-COCO and Open Images subsets, and the CC3M portion uses only image-caption pairs whose images were still valid in January 2024.
- The updated 100M-word text-only BabyLM corpus drops the QED portion of the 2023 dataset in favour of a larger CHILDES share of 29M words. The remaining sources keep their previous relative proportions: 26M words of children's Project Gutenberg text, 20M of OpenSubtitles, 15M of Simple English Wikipedia, 8M of BNC dialogue and 1M of Switchboard. (Table 1)
Scope: Strict-track word counts, which are approximate; QED was replaced because on inspection its quality was poorer than hoped.
- The 2nd BabyLM Challenge baselines are built from the previous year's winning submissions rather than trained naively. The Strict and Strict-small baselines are GPT-2, LTG-BERT and Contextualizer, and the Vision baselines are GIT and Flamingo. (Section 5.1)
Scope: Baselines as announced in the call for papers; final released checkpoints and their scores are not reported in the call.
- The BabyLM Challenge places no limit on the number of training epochs or on hyperparameters. The organizers report that in their internal results over-training beyond a couple of epochs gives minor gains at most. (Section 7)
Scope: The internal-results statement is an unquantified organizer observation; the stated rationale also covers engineering practice at these data scales and human re-access to linguistic memories.
- Submissions to the BabyLM Challenge must be able to assign a (pseudo) log-likelihood to a string of text and to be fine-tuned for classification, but need not be able to generate sequences. (Section 5)
Scope: The 2024 evaluation pipeline is built on catwalk so that models outside the HuggingFace transformers library can be submitted; Vision-track models must score text conditioned on an image.
- BabyLM runs its own review process with acceptance based on soundness alone, planning to reject only submissions that make incorrect or unjustified claims. Papers are archival and may be up to 8 pages, with dual submission allowed but not dual publication. (Section 6.3)
Scope: Review policy as stated for the 2024/2025 edition; the presentation venue and formatting requirements were not yet finalized at the time of the call.
- The BabyLM organizers report that submissions to the first challenge did not gain from non-linguistic grounding, and invite the 2024/2025 vision-language track partly to revisit that question. (Section 7)
Scope: A summary statement about the 2023 edition's submissions, concerning non-linguistic grounding signals rather than multimodal training in general.
- The BabyLM Challenge is a recurring shared task that reorients pretraining research toward sample efficiency by capping training data at a developmentally plausible 10M or 100M words. A stated goal is making pretraining research feasible on a university budget. (Section 1)
Scope: A shared-task call for papers, not an empirical study; the rationale is argued at greater length in the 2023 call and proceedings introduction by Warstadt et al. Second edition, run in 2024/2025.
- The 2nd BabyLM Challenge's paper-only track opens the shared task to non-model submissions, such as novel cognitively-inspired evaluation metrics and in-depth analyses of individual BabyLM models. Best-paper selection in that track is not driven by evaluation scores. (Section 3)
Scope: The paper track replaces the loose track used in the 2023 edition; papers may still include a model and report its scores.
Common misreadings
- The 2nd BabyLM Challenge does not require using the organizers' corpus: participants may build any dataset they like as long as it stays within the 100M-word or 10M-word budget and comes with a datasheet.
- The 100M-word cap in the BabyLM Challenge is not a cap on tokens seen during training. Multiple epochs and augmentation are allowed; what is capped is the total amount of distinct text any component of the system learned from.
- The Vision track of the 2nd BabyLM Challenge is not evaluated on multimodal tasks alone — its submissions are also run on the language-only evaluation suite.
- The organizers' remark that the first challenge's submissions did not benefit from non-linguistic grounding is not a conclusion that grounding cannot help; the vision-language track exists to encourage further work on exactly that question.
- The Strict-small track's 10M-word budget is a separate track, not a smaller warm-up for Strict; a paper may enter several tracks with different models.
Terminology in this paper
- Strict track
- BabyLM Challenge track for language models pretrained on 100M words of text or less, evaluated on language-only tasks.
- Strict-small track
- BabyLM Challenge track for language models pretrained on 10M words of text or less, evaluated on language-only tasks.
- Vision track
- BabyLM Challenge track for multimodal image-text models trained within a 100M word budget, which must be able to assign (pseudo) log-likelihoods to text conditioned on an image and are evaluated on both language-only and multimodal tasks.
- Paper track
- BabyLM Challenge track for submissions that need not include a competition model, such as new cognitively-inspired evaluation metrics or analyses of an existing BabyLM model.
- Developmentally plausible corpus
- a pretraining corpus whose size and composition approximate the language input a human child receives, on the order of 10M to 100M words drawn from sources such as child-directed speech, dialogue, subtitles and children's books.
References
See the full reference list in the paper.