# User Feedback Provides a Unique Signal that LLMs Can not Detect Authors: Shachar Don-Yehiya, Leshem Choshen, Omri Abend Venue: preprint (2026) ## What this paper shows Revising an LLM answer with the user's feedback fixes far more errors than revising it without, yet LLM judges often prefer the unfixed revision, most of all on errors the judging model cannot fix alone, so standard pairwise evaluation makes useful feedback look useless. ## Claims, with scope - With user feedback that says how to fix a corrupted answer, Gemini-3-Flash resolves the error in 95.6% of cases versus 86.3% without it. Qwen3-8B rises from 49.2% to 80.8%. Scope: 500 Arena-Hard-v2.0 hard-prompt queries with o3 responses, each corrupted 4 ways; resolution scored by a Gemini-3.1-Pro-preview evaluator. Evidence: Table 4 - On real user feedback from ShareLM conversations, feedback-informed revisions resolve the issue in 89% of cases for Gemini-3-Flash versus 70% without, and in 53% versus 26% for Qwen3-8B. Scope: 1000 English ShareLM feedback instances kept as generally applicable; the LLM resolution evaluator agreed with a human at Cohen's kappa of only 0.24 here. Evidence: Section 5.1, Figure 1 - A pairwise LLM judge prefers Gemini-3-Flash's without-feedback revision 37.8% of the time and its with-feedback revision only 27.2%, although the with-feedback revision fixes more corruptions. For Qwen3-8B revisions the preference reverses, 51.2% versus 34.3%. Scope: Gemini-3.1-Pro-preview judge, random response order, averaged over 4 corruption types on Arena-Hard-v2.0 hard prompts; 35.0% of Gemini-3-Flash comparisons were ties. Evidence: Table 1 - On real ShareLM user feedback, a Gemini-3.1-Pro judge prefers the without-feedback revision for both improvers, 55.4% versus 32.5% for Gemini-3-Flash and 51.8% versus 35.0% for Qwen3-8B. Scope: 1000 English ShareLM samples; ties were 12.1% and 13.2%, and no known ground truth exists for these real conversations. Evidence: Table 7 - When only the with-feedback revision fixed a corrupted answer, a Gemini-3.1-Pro judge still picks it in only 54.6% of Gemini-3-Flash cases, against 80.8% for Qwen3-8B revisions. Scope: WFShouldWin subsets of 205 (Gemini-3-Flash) and 574 (Qwen3-8B) Arena-Hard-v2.0 hard-prompt samples, where the with-feedback revision is the only valid one. Evidence: Table 11 (Section 5.3) - On real user feedback where only the with-feedback revision resolved the issue, a Gemini-3.1-Pro judge prefers it in only 34.0% of Gemini-3-Flash cases and 54.0% of Qwen3-8B cases. Scope: 215 (Gemini-3-Flash) and 274 (Qwen3-8B) English ShareLM samples; resolution labels come from an LLM evaluator. Evidence: Table 8 (Section 5.3) - Gemini-3-Flash as a judge picks the correct feedback-fixed response 87.7% of the time on queries it could fix alone, but 72.2% on queries it fixed only with feedback, 15.5 points lower. Scope: Qwen3-8B revisions on its synthetic WFShouldWin subset, so the correct answer is known; queries split by whether Gemini-3-Flash resolved them without feedback. Evidence: Table 12 (Section 5.4) - Qwen3-8B judging its own revisions picks the correct with-feedback response in 67.6% of cases, 13 points below the 80.8% of a Gemini-3.1-Pro judge. For Gemini-3-Flash, self-judge and Gemini-3.1-Pro agree at 54.4% and 54.6%. Scope: Synthetic hard-prompt WFShouldWin subsets of each improver; the self-judge is the same model that wrote both revisions. Evidence: Figure 3 (Section 6.1) - Revisions made with feedback contain a larger share of content fixes (factuality, completeness, logic) relative to style edits, and fewer no-improvement cases, than revisions made without feedback. Scope: Gemini-3-Flash revisions on its WFShouldWin subset of the full synthetic data including creative writing; one primary type per response, assigned by an LLM classifier. Evidence: Figure 4 (Section 6.2) - Feedback that only names the problem still raises issue resolution by 6.3 points for Gemini-3-Flash and 16.7 for Qwen3-8B, and feedback saying only that something is wrong adds 3.4 and 10.3. Scope: 100 queries per corruption type from Arena-Hard-v2.0 hard prompts; the one exception is causality inversion, where the bare-wrong signal lowered Gemini-3-Flash resolution by 5.2 points. Evidence: Table 2 (Section 6.3), Table 10 - When a second, feedback-based revision fixes a response that a second feedback-free revision leaves broken, a Gemini-3.1-Pro judge prefers the fixed one in only 34.2% of cases. Scope: Gemini-3-Flash improver on its synthetic WFShouldWin subset, n=123 such cases; across all pairs the judge favours the twice-without-feedback revision 57.3% to 28.9%. Evidence: Appendix F, Table 13 - With GPT-OSS-20B as the improver, feedback raises issue resolution from 48.5% to 54.7%, and the pairwise judge picks the with-feedback revision in 95.2% of cases where only it fixed the error. Scope: Initial results on about 400 synthetic hard-prompt samples, averaged across the 4 corruption types. Evidence: Tables 14 and 15 (Appendix G) - Shown an original o3 response beside its minimally corrupted version, a Gemini-3.1-Pro judge prefers the original in 96.6% to 99.4% of cases across the 4 corruption types. Scope: Arena-Hard-v2.0 synthetic data; a direct original-versus-corrupted comparison, not a comparison between two revisions. Evidence: Table 3 (Appendix A) - "User Feedback Provides a Unique Signal that LLMs Can not Detect" argues that naturally occurring user feedback is a strong improvement signal whose value is masked by LLM-judge evaluation. It tests this with ground-truth corruptions and real user feedback. Scope: As of the September 2026 arXiv preprint; test-time revision only, with Gemini-3-Flash and Qwen3-8B improvers and a Gemini-3.1-Pro judge, set against prior findings that such feedback is noisy. - Don-Yehiya, Choshen and Abend propose that because LLM judges fail to credit feedback-driven fixes, user feedback carries information not encoded in model weights, which distillation or self-improvement could not supply. Scope: Argued from inference-only revision experiments; no distillation, self-improvement or training on feedback was run. ## Common misreadings - LLM judges do not always prefer the feedback-free revision: for Qwen3-8B revisions on synthetic data the Gemini-3.1-Pro judge prefers the with-feedback revision 51.2% to 34.3%, and for GPT-OSS-20B it picks the only correct revision 95.2% of the time. - The judge failure is not an inability to spot errors at all. Comparing an original o3 response with its corrupted version, the Gemini-3.1-Pro judge prefers the original in 96.6% to 99.4% of cases. - The naturalistic issue-resolution rates are an approximation: the LLM evaluator agreed with a human annotator at Cohen's kappa of 0.24 on real user data, against 0.81 on the synthetic data. - The main synthetic results cover Arena-Hard-v2.0 hard prompts only, because corruptions were valid in just 59.4% of sampled creative-writing items versus 92.2% of hard-prompt items. - All experiments revise responses at test time. No model was trained on user feedback, and the idea that feedback carries information absent from model weights is argued, not measured. ## Terminology - WFShouldWin: The subset of queries on which the revision made with feedback resolves the issue while the revision made without feedback does not, so a correct pairwise judge must prefer the with-feedback revision. - self-improvable subset: Queries on which an improver model resolves a corrupted response both with and without feedback. - feedback-improvable subset: Queries on which an improver model resolves a corrupted response only when given feedback, equivalent to that model's WFShouldWin set. - issue resolution evaluation: An LLM evaluator's binary verdict on whether a revised response addresses the fix-it feedback for a known error, given the query, the corrupted response and the feedback. - response corruption: Deliberately inserting one minimal error into a valid model response, by causality inversion, crucial omission, entity or subject swap, or logic operator reversal, so a ground-truth fix exists. ## Links - arXiv: https://arxiv.org/abs/2609.02859 - PDF: https://arxiv.org/pdf/2609.02859 - HTML: https://arxiv.org/html/2609.02859 - Hugging Face: https://huggingface.co/papers/2609.02859 - alphaXiv: https://www.alphaxiv.org/abs/2609.02859 - Semantic Scholar: https://www.semanticscholar.org/paper/291615425 ## How to cite @misc{donyehiya2026user, title = {User Feedback Provides a Unique Signal that LLMs Can not Detect}, author = {Shachar Don-Yehiya and Leshem Choshen and Omri Abend}, year = {2026}, }