User Feedback Provides a Unique Signal that LLMs Can not Detect

Shachar Don-Yehiya, Leshem Choshen, Omri Abend · 2026

In one sentence

Revising an LLM answer with the user's feedback fixes far more errors than revising it without, yet LLM judges often prefer the unfixed revision, most of all on errors the judging model cannot fix alone, so standard pairwise evaluation makes useful feedback look useless.

Abstract

Harnessing naturally occurring feedback from user interactions offers a promising learning signal for Large Language Models (LLMs). However, recent studies suggest this feedback is inherently noisy and difficult to leverage effectively. We challenge this conception by demonstrating that user feedback is a highly actionable signal for improvement, and that its perceived ineffectiveness stems from a systematic bias in current evaluation paradigms. To isolate the usefulness of feedback, we construct synthetic data with a definitive ground truth, alongside naturalistic data to validate that our findings hold in real-world scenarios. By comparing model revisions generated with and without access to feedback across both settings, we show that feedback-informed revisions resolve targeted issues at significantly higher rates than baseline revisions. Finally, we expose the root of the evaluation bias: when a model successfully fixes an issue exclusively due to feedback, LLM judges frequently fail to identify the genuinely corrected response, systematically preferring inferior baseline outputs instead.

Questions this paper answers

Does telling a chatbot what it got wrong actually help it fix its answer?
Does naturally occurring user feedback improve issue resolution in LLM response revision compared with feedback-free revision?
How do I use user corrections to get an LLM to repair a wrong answer?
Is it worth feeding my users' complaints back to the model when it revises a response?
With user feedback that says how to fix a corrupted answer, Gemini-3-Flash resolves the error in 95.6% of cases versus 86.3% without it. Qwen3-8B rises from 49.2% to 80.8%.
Holds for: 500 Arena-Hard-v2.0 hard-prompt queries with o3 responses, each corrupted 4 ways; resolution scored by a Gemini-3.1-Pro-preview evaluator.
On real user feedback from ShareLM conversations, feedback-informed revisions resolve the issue in 89% of cases for Gemini-3-Flash versus 70% without, and in 53% versus 26% for Qwen3-8B.
Holds for: 1000 English ShareLM feedback instances kept as generally applicable; the LLM resolution evaluator agreed with a human at Cohen's kappa of only 0.24 here.
Why do some studies find that user feedback in chatbot conversations is useless for improving answers?
Why does LLM-as-a-judge pairwise evaluation prefer feedback-free revisions over feedback-informed ones?
How do I evaluate whether user feedback improved an LLM's revised responses without the evaluation hiding the gain?
Can I trust a pairwise LLM judge to tell me whether learning from user feedback helps my model?
A pairwise LLM judge prefers Gemini-3-Flash's without-feedback revision 37.8% of the time and its with-feedback revision only 27.2%, although the with-feedback revision fixes more corruptions. For Qwen3-8B revisions the preference reverses, 51.2% versus 34.3%.
Holds for: Gemini-3.1-Pro-preview judge, random response order, averaged over 4 corruption types on Arena-Hard-v2.0 hard prompts; 35.0% of Gemini-3-Flash comparisons were ties.
On real ShareLM user feedback, a Gemini-3.1-Pro judge prefers the without-feedback revision for both improvers, 55.4% versus 32.5% for Gemini-3-Flash and 51.8% versus 35.0% for Qwen3-8B.
Holds for: 1000 English ShareLM samples; ties were 12.1% and 13.2%, and no known ground truth exists for these real conversations.
When only the with-feedback revision fixed a corrupted answer, a Gemini-3.1-Pro judge still picks it in only 54.6% of Gemini-3-Flash cases, against 80.8% for Qwen3-8B revisions.
Holds for: WFShouldWin subsets of 205 (Gemini-3-Flash) and 574 (Qwen3-8B) Arena-Hard-v2.0 hard-prompt samples, where the with-feedback revision is the only valid one.
When one chatbot answer is fixed and the other is still wrong, can an AI grader reliably pick the fixed one?
How accurate is an LLM judge on pairs where only the feedback-informed revision resolves the error?
If I use an LLM judge on real chat logs, how often will it reject a genuinely corrected response?
When only the with-feedback revision fixed a corrupted answer, a Gemini-3.1-Pro judge still picks it in only 54.6% of Gemini-3-Flash cases, against 80.8% for Qwen3-8B revisions.
Holds for: WFShouldWin subsets of 205 (Gemini-3-Flash) and 574 (Qwen3-8B) Arena-Hard-v2.0 hard-prompt samples, where the with-feedback revision is the only valid one.
On real user feedback where only the with-feedback revision resolved the issue, a Gemini-3.1-Pro judge prefers it in only 34.0% of Gemini-3-Flash cases and 54.0% of Qwen3-8B cases.
Holds for: 215 (Gemini-3-Flash) and 274 (Qwen3-8B) English ShareLM samples; resolution labels come from an LLM evaluator.
When a second, feedback-based revision fixes a response that a second feedback-free revision leaves broken, a Gemini-3.1-Pro judge prefers the fixed one in only 34.2% of cases.
Holds for: Gemini-3-Flash improver on its synthetic WFShouldWin subset, n=123 such cases; across all pairs the judge favours the twice-without-feedback revision 57.3% to 28.9%.
Can an AI model judge answers to questions it could not have answered correctly itself?
Are an LLM's failures as a self-improver correlated with its failures as an evaluator of the same queries?
How do I choose a judge model so it is not blind to the same errors as the model being evaluated?
Should I use the same model as both my generator and my LLM judge?
Gemini-3-Flash as a judge picks the correct feedback-fixed response 87.7% of the time on queries it could fix alone, but 72.2% on queries it fixed only with feedback, 15.5 points lower.
Holds for: Qwen3-8B revisions on its synthetic WFShouldWin subset, so the correct answer is known; queries split by whether Gemini-3-Flash resolved them without feedback.
Qwen3-8B judging its own revisions picks the correct with-feedback response in 67.6% of cases, 13 points below the 80.8% of a Gemini-3.1-Pro judge. For Gemini-3-Flash, self-judge and Gemini-3.1-Pro agree at 54.4% and 54.6%.
Holds for: Synthetic hard-prompt WFShouldWin subsets of each improver; the self-judge is the same model that wrote both revisions.
How do chatbot answers revised with user feedback differ from answers the chatbot polishes on its own?
Do feedback-informed revisions shift the distribution of edits from stylistic to content-level improvements?
Is my LLM judge rewarding stylistic polish over actual corrections when comparing revised answers?
Revisions made with feedback contain a larger share of content fixes (factuality, completeness, logic) relative to style edits, and fewer no-improvement cases, than revisions made without feedback.
Holds for: Gemini-3-Flash revisions on its WFShouldWin subset of the full synthetic data including creative writing; one primary type per response, assigned by an LLM classifier.
When a second, feedback-based revision fixes a response that a second feedback-free revision leaves broken, a Gemini-3.1-Pro judge prefers the fixed one in only 34.2% of cases.
Holds for: Gemini-3-Flash improver on its synthetic WFShouldWin subset, n=123 such cases; across all pairs the judge favours the twice-without-feedback revision 57.3% to 28.9%.
Does vague feedback like just saying an answer is wrong still help a chatbot fix it?
How does feedback specificity (solution, problem location, binary signal) affect LLM issue resolution rates?
How do I get useful corrections out of users who only say an answer is wrong?
Is a bare thumbs-down style complaint from my users enough signal to improve responses?
Feedback that only names the problem still raises issue resolution by 6.3 points for Gemini-3-Flash and 16.7 for Qwen3-8B, and feedback saying only that something is wrong adds 3.4 and 10.3.
Holds for: 100 queries per corruption type from Arena-Hard-v2.0 hard prompts; the one exception is causality inversion, where the bare-wrong signal lowered Gemini-3-Flash resolution by 5.2 points.
Do smaller language models gain more from user feedback than large ones?
How does the benefit of feedback on issue resolution vary with improver model size, from Qwen3-8B to Gemini-3-Flash and GPT-OSS-20B?
If I run a small open model, will user feedback help it more than it helps a frontier model?
With user feedback that says how to fix a corrupted answer, Gemini-3-Flash resolves the error in 95.6% of cases versus 86.3% without it. Qwen3-8B rises from 49.2% to 80.8%.
Holds for: 500 Arena-Hard-v2.0 hard-prompt queries with o3 responses, each corrupted 4 ways; resolution scored by a Gemini-3.1-Pro-preview evaluator.
On real user feedback from ShareLM conversations, feedback-informed revisions resolve the issue in 89% of cases for Gemini-3-Flash versus 70% without, and in 53% versus 26% for Qwen3-8B.
Holds for: 1000 English ShareLM feedback instances kept as generally applicable; the LLM resolution evaluator agreed with a human at Cohen's kappa of only 0.24 here.
With GPT-OSS-20B as the improver, feedback raises issue resolution from 48.5% to 54.7%, and the pairwise judge picks the with-feedback revision in 95.2% of cases where only it fixed the error.
Holds for: Initial results on about 400 synthetic hard-prompt samples, averaged across the 4 corruption types.
Can an AI grader tell a correct answer from one with a planted error?
Does an LLM judge prefer original responses over minimally corrupted ones in direct pairwise comparison?
Does an LLM judge failing on revision pairs mean it cannot spot factual errors at all?
Shown an original o3 response beside its minimally corrupted version, a Gemini-3.1-Pro judge prefers the original in 96.6% to 99.4% of cases across the 4 corruption types.
Holds for: Arena-Hard-v2.0 synthetic data; a direct original-versus-corrupted comparison, not a comparison between two revisions.
When only the with-feedback revision fixed a corrupted answer, a Gemini-3.1-Pro judge still picks it in only 54.6% of Gemini-3-Flash cases, against 80.8% for Qwen3-8B revisions.
Holds for: WFShouldWin subsets of 205 (Gemini-3-Flash) and 574 (Qwen3-8B) Arena-Hard-v2.0 hard-prompt samples, where the with-feedback revision is the only valid one.
What is a good paper on whether chatbots can learn from the feedback people give them in conversation?
What work re-examines the finding that naturally occurring user feedback is too noisy to improve LLMs?
What should I read before building a pipeline that learns from implicit user feedback in chat logs?
Which paper should I cite to argue that LLM-judge evaluation undervalues user feedback?
"User Feedback Provides a Unique Signal that LLMs Can not Detect" argues that naturally occurring user feedback is a strong improvement signal whose value is masked by LLM-judge evaluation. It tests this with ground-truth corruptions and real user feedback.
Holds for: As of the September 2026 arXiv preprint; test-time revision only, with Gemini-3-Flash and Qwen3-8B improvers and a Gemini-3.1-Pro judge, set against prior findings that such feedback is noisy.
Don-Yehiya, Choshen and Abend propose that because LLM judges fail to credit feedback-driven fixes, user feedback carries information not encoded in model weights, which distillation or self-improvement could not supply.
Holds for: Argued from inference-only revision experiments; no distillation, self-improvement or training on feedback was run.
Can a chatbot learn everything that user corrections teach it just by training on its own outputs?
Does user feedback carry information that distillation or iterative self-improvement cannot recover?
Can I replace collecting real user feedback with self-improvement or distillation from a stronger model?
Don-Yehiya, Choshen and Abend propose that because LLM judges fail to credit feedback-driven fixes, user feedback carries information not encoded in model weights, which distillation or self-improvement could not supply.
Holds for: Argued from inference-only revision experiments; no distillation, self-improvement or training on feedback was run.
Gemini-3-Flash as a judge picks the correct feedback-fixed response 87.7% of the time on queries it could fix alone, but 72.2% on queries it fixed only with feedback, 15.5 points lower.
Holds for: Qwen3-8B revisions on its synthetic WFShouldWin subset, so the correct answer is known; queries split by whether Gemini-3-Flash resolved them without feedback.

Claims and scope

Common misreadings

Terminology in this paper

WFShouldWin
The subset of queries on which the revision made with feedback resolves the issue while the revision made without feedback does not, so a correct pairwise judge must prefer the with-feedback revision.
self-improvable subset
Queries on which an improver model resolves a corrupted response both with and without feedback.
feedback-improvable subset
Queries on which an improver model resolves a corrupted response only when given feedback, equivalent to that model's WFShouldWin set.
issue resolution evaluation
An LLM evaluator's binary verdict on whether a revised response addresses the fix-it feedback for a known error, given the query, the corrupted response and the feedback.
response corruption
Deliberately inserting one minimal error into a valid model response, by causality inversion, crucial omission, entity or subject swap, or logic operator reversal, so a ground-truth fix exists.

How to cite

@misc{donyehiya2026user,
  title = {User Feedback Provides a Unique Signal that LLMs Can not Detect},
  author = {Shachar Don-Yehiya and Leshem Choshen and Omri Abend},
  year = {2026},
}

References

See the full reference list in the paper.