Papers
- TIES-Merging: Resolving Interference When Merging Models Advances in Neural Information Processing Systems 36: Annual · 855 citations
- tinyBenchmarks: evaluating LLMs with fewer examples Forty-first International Conference on Machine Learning, {I · 276 citations
- Active Learning for BERT: An Empirical Study Proceedings of the 2020 Conference on Empirical Methods in N · 244 citations
- Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora Proceedings of the BabyLM Challenge at the 27th Conference o · 232 citations
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation Proceedings of the 63rd Annual Meeting of the Association fo · 181 citations
- An autonomous debating system Nat. · 172 citations
- Q²: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering CoRR · 167 citations
- On the Weaknesses of Reinforcement Learning for Neural Machine Translation 8th International Conference on Learning Representations, {I · 127 citations
- Fusing finetuned models for better pretraining ArXiv · 120 citations
- DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question Answering Proceedings of the 61st Annual Meeting of the Association fo · 119 citations
- Model merging with SVD to tie the Knots International Conference on Learning Representations · 111 citations
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty The Fourteenth International Conference on Learning Represen · 97 citations
- Jump to Conclusions: Short-Cutting Transformers with Linear Transformations Proceedings of the 2024 Joint International Conference on Co · 97 citations
- Asymmetry in Low-Rank Adapters of Foundation Models Forty-first International Conference on Machine Learning, {I · 90 citations
- Efficient multi-prompt evaluation of LLMs The Thirty-eighth Annual Conference on Neural Information Pr · 83 citations
- Call for Papers - The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus CoRR · 82 citations
- Are You Convinced? Choosing the More Convincing Evidence with a Siamese Network Proceedings of the 57th Conference of the Association for Co · 77 citations
- Will it Blend? Blending Weak and Strong Labeled Data in a Neural Network for Argumentation Mining Proceedings of the 56th Annual Meeting of the Association fo · 72 citations
- Knowledge is a Region in Weight Space for Fine-tuned Language Models Findings of the Association for Computational Linguistics: { · 71 citations
- Corpus Wide Argument Mining - A Working Solution The Thirty-Fourth {AAAI} Conference on Artificial Intelligen · 70 citations
- DORA The Explorer: Directed Outreaching Reinforcement Action-Selection 6th International Conference on Learning Representations, {I · 70 citations
- A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning Transactions on Machine Learning Research · 68 citations
- Efficient Benchmarking (of Language Models) Proceedings of the 2024 Conference of the North American Cha · 68 citations
- Elements of World Knowledge (EWOK): A cognition-inspired framework for evaluating basic world knowledge in language models CoRR · 67 citations
- ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning Proceedings of the 61st Annual Meeting of the Association fo · 64 citations
- Let's Agree to Agree: Neural Networks Share Classification Order on Real Datasets Proceedings of the 37th International Conference on Machine · 64 citations
- Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora The 2nd BabyLM Challenge at the 28th Conference on Computati · 62 citations
- NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning Proceedings of the 2024 Conference on Empirical Methods in N · 46 citations
- The Grammar-Learning Trajectories of Neural Language Models Proceedings of the 60th Annual Meeting of the Association fo · 41 citations
- [Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus CoRR · 39 citations
- SemEval-2019 Task 1: Cross-lingual Semantic Parsing with UCCA Proceedings of the 13th International Workshop on Semantic E · 38 citations
- Inherent Biases in Reference-based Evaluation for Grammatical Error Correction Proceedings of the 56th Annual Meeting of the Association fo · 38 citations
- Genie: Achieving Human Parity in Content-Grounded Datasets Generation CoRR · 37 citations
- Reference-less Measure of Faithfulness for Grammatical Error Correction Proceedings of the 2018 Conference of the North American Cha · 36 citations
- Automatic Metric Validation for Grammatical Error Correction Proceedings of the 56th Annual Meeting of the Association fo · 35 citations
- BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop arXiv.org · 33 citations
- Learning to combine Grammatical Error Corrections Proceedings of the Fourteenth Workshop on Innovative Use of · 31 citations
- Bigger is not always better: The importance of human-scale language modeling for psycholinguistics Journal of Memory and Language · 30 citations
- Where to start? Analyzing the potential value of intermediate models Proceedings of the 2023 Conference on Empirical Methods in N · 30 citations
- Cluster & Tune: Boost Cold Start Performance in Text Classification Proceedings of the 60th Annual Meeting of the Association fo · 28 citations
- Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead Forty-second International Conference on Machine Learning · 27 citations
- Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability Findings of the Association for Computational Linguistics, { · 27 citations
- Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families The Thirty-ninth Annual Conference on Neural Information Pro · 26 citations
- The Language of Legal and Illegal Activity on the Darknet Proceedings of the 57th Conference of the Association for Co · 25 citations
- Data Contamination Report from the 2024 CONDA Shared Task CoRR · 23 citations
- Classifying Syntactic Errors in Learner Language Proceedings of the 24th Conference on Computational Natural · 23 citations
- A Hitchhiker's Guide to Scaling Law Estimation Forty-second International Conference on Machine Learning · 22 citations
- Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI Proceedings of the 2024 Conference of the North American Cha · 22 citations
- DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation Findings of the Association for Computational Linguistics: A · 21 citations
- ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization Transactions on Machine Learning Research · 21 citations
- Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench preprint · 21 citations
- Label Sleuth: From Unlabeled Text to a Classifier in a Few Hours Proceedings of the The 2022 Conference on Empirical Methods · 21 citations
- ZipNN: Lossless Compression for AI Models 2025 IEEE 18th International Conference on Cloud Computing ( · 20 citations
- Automatically Extracting Challenge Sets for Non-Local Phenomena in Neural Machine Translation Proceedings of the 23rd Conference on Computational Natural · 20 citations
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation preprint · 19 citations
- LiveXiv - A Multi-Modal live benchmark based on Arxiv papers content The Thirteenth International Conference on Learning Represen · 18 citations
- Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs NAACL · 16 citations
- Label-Efficient Model Selection for Text Generation Proceedings of the 62nd Annual Meeting of the Association fo · 16 citations
- Human Learning by Model Feedback: The Dynamics of Iterative Prompting with Midjourney Proceedings of the 2023 Conference on Empirical Methods in N · 16 citations
- Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion CoRR · 16 citations
- Mediators in Determining what Processing BERT Performs First Proceedings of the 2021 Conference of the North American Cha · 16 citations
- The Future of Open Human Feedback Nature Machine Intelligence · 15 citations
- Lossless and Near-Lossless Compression for Foundation Models CoRR · 15 citations
- Unsupervised Expressive Rules Provide Explainability and Assist Human Experts Grasping New Domains Findings of the Association for Computational Linguistics: { · 15 citations
- Naturally Occurring Feedback is Common, Extractable and Useful preprint · 14 citations
- The Mighty ToRR: A Benchmark for Table Reasoning and Robustness arXiv · 13 citations
- Semantics-aware Attention Improves Neural Machine Translation Proceedings of the 11th Joint Conference on Lexical and Comp · 13 citations
- Findings of the Third BabyLM Challenge: Accelerating Language Modeling Research with Cognitively Plausible Data Proceedings of the First BabyLM Workshop · 12 citations
- PreQuEL: Quality Estimation of Machine Translation Outputs in Advance Proceedings of the 2022 Conference on Empirical Methods in N · 12 citations
- SERRANT: a syntactic classifier for English Grammatical Error Types CoRR · 12 citations
- The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community Proceedings of the 63rd Annual Meeting of the Association fo · 11 citations
- Benchmark Agreement Testing Done Right: A Guide for LLM Benchmark Evaluation CoRR · 11 citations
- Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures preprint · 10 citations
- Neurips 2023 llm efficiency fine-tuning competition arXiv preprint arXiv:2503.13507 · 8 citations
- Navigating the Modern Evaluation Landscape: Considerations in Benchmarks and Frameworks for Large Language Models (LLMs) Proceedings of the 2024 Joint International Conference on Co · 8 citations
- Reinforcement Learning with Large Action Spaces for Neural Machine Translation Proceedings of the 29th International Conference on Computat · 8 citations
- GrASP: A Library for Extracting and Exploring Human-Interpretable Textual Patterns Proceedings of the Thirteenth Language Resources and Evaluat · 8 citations
- Transition based Graph Decoder for Neural Machine Translation CoRR · 8 citations
- CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data arXiv preprint arXiv:2601.18026 · 7 citations
- TextArena arXiv preprint arXiv:2504.11442 · 7 citations
- Holmes: Benchmark the Linguistic Competence of Language Models CoRR (accepted to TaCL) · 7 citations
- General Agent Evaluation preprint · 6 citations
- ComSum: Commit Messages Summarization and Meaning Preservation CoRR · 6 citations
- ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models arXiv preprint arXiv:2601.15812 · 5 citations
- Pretraining Language Models for Diachronic Linguistic Change Discovery EACL · 5 citations
- Do LLMs Benefit From Their Own Words? arXiv.org · 5 citations
- BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data EACL · 4 citations
- Unforgettable Generalization in Language Models First Conference on Language Modeling · 4 citations
- MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs arXiv.org · 3 citations
- CUBE: A Standard for Unifying Agent Benchmarks arXiv.org · 3 citations
- From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation Conference on Empirical Methods in Natural Language Processi · 3 citations
- Robustness as an Emergent Property of Task Performance arXiv preprint arXiv:2602.03344 · 2 citations
- Will it Merge? On The Causes of Model Mergeability arXiv preprint arXiv:2601.06672 · 2 citations
- Mediocrity is the key for LLM as a Judge Anchor Selection Annual Meeting of the Association for Computational Linguist · 2 citations
- LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users arXiv preprint arXiv:2507.02850 · 2 citations
- On Neurons Invariant to Sentence Structural Changes in Neural Machine Translation Proceedings of the 26th Conference on Computational Natural · 2 citations
- SemEval 2019 Shared Task: Cross-lingual Semantic Parsing with UCCA - Call for Participation arXiv.org · 2 citations
- BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop arXiv.org · 1 citations
- Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results preprint · 1 citations
- Automated Discovery Has No Universally Superior Harness preprint · 1 citations
- How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability preprint · 1 citations
- A Latent Variable Framework for Scaling Laws in Large Language Models preprint · 1 citations
- Insights from the first BabyLM Challenge: Training sample-efficient language models on a developmentally plausible corpus Proceedings of the Annual Meeting of the Cognitive Science S · 1 citations
- Enhancing the Transformer Decoder with Transition-based Syntax Proceedings of the 26th Conference on Computational Natural · 1 citations
- Part of Speech and Universal Dependency effects on English Arabic Machine Translation CoRR · 1 citations
- All Neural Networks are Created Equal arXiv.org · 1 citations
- Position: Agentic Systems Should be General preprint
- Cross-Lingual Exploration for Parametric Knowledge preprint
- Instructions Shape Production of Language, not Processing arXiv.org
- Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting preprint
- Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data Annual Meeting of the Association for Computational Linguist
- Resolving Interference (RI): Disentangling Models for Improved Model Merging arXiv.org
- CRISP: Complex Reasoning with Interpretable Step-based Plans arXiv preprint arXiv:2507.08037
- Can Gradient Descent Simulate Prompting? arXiv preprint arXiv:2506.20989
- Tie the KnOTS: Model Merging with SVD ICLR
- LLM Merging: Building LLMs Efficiently through Merging NeurIPS 2024 Competition Track
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging arXiv preprint arXiv:2402.02622
- High-dimensional Learning Dynamics 2024: The Emergence of Structure and Reasoning Workshop at ICML24 ICML
- Super Tiny Language Models arXiv preprint arXiv:2405.14159
- A framework for few-shot language model evaluation Zenodo
- The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models Proceedings of the 62nd Annual Meeting of the Association fo
- Can You Trust Your Metric? Automatic Concatenation-Based Tests for Metric Validity arXiv preprint arXiv:2408.12259
- The llama 3 herd of models arXiv preprint arXiv:2407.21783
- Not all layers are equally as important: Every Layer Counts BERT arXiv preprint arXiv:2311.02265
- MuLER: Detailed and Scalable Reference-based Evaluation Proceedings of the 27th Conference on Computational Natural
- Resolving Interference When Merging Models CoRR
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time ICML
- Some Grammatical Errors are Frequent, Others are Important CoRR
- Holistic evaluation of language models arXiv preprint arXiv:2211.09110
- Merging models with fisher-weighted averaging arXiv:2111.09832
- Inherent Biases in Reference-based Evaluation for Grammatical Error Correction and Text Simplification CoRR
- Attention is all you need Advances in Neural Information Processing Systems
- Mapping the early language environment using all-day recordings and automated analysis American journal of speech-language pathology
- Sapiens: A brief history of humankind Random House
- European public acceptance of euthanasia: socio-demographic and cultural factors associated with the acceptance of euthanasia in 33 European countries Social science \& medicine