Do LLMs Benefit From Their Own Words?
when a model's own earlier response carries errors, hallucinations or style into later turns
Jenny Y. Huang, Leshem Choshen, Ramon Astudillo, Tamara Broderick, Jacob Andreas · 2026
In one sentence
Replacing or dropping prior assistant responses in multi-turn chat histories from WildChat and ShareLM often matches full-context response quality while using roughly 8x less context, because 36.4% of in-the-wild turns are self-contained and past responses can pollute later ones.
Abstract
In multi-turn conversations, large language models typically condition on the full conversation history: both past user prompts and assistant responses. We revisit this design choice by comparing full-context prompting to four alternative, substantially-reduced context configurations. Analyzing in-the-wild multi-turn conversations across three open reasoning and one state-of-the-art model, we find that response quality is largely preserved under aggressive context filtering: replacing all prior assistant turns with one-sentence summaries or keeping only the most recent user--assistant exchange often matches storing full context in performance while using roughly 8x less context. To understand this result, we observe that a substantial fraction of user turns (36.4%) in multi-turn conversations are self-contained and that many follow-up turns can be addressed by seeing only the immediately preceding user--assistant exchange. Furthermore, we find that when models condition on their own past responses, this can lead to context pollution, a phenomenon in which reasoning errors, hallucinations, or stylistic artifacts propagate across turns. Motivated by these findings, we design a context-filtering approach that selectively omits the assistant-side history. Taken together, these findings suggest moving away from storing full dialogue transcripts and instead retaining only what is relevant.
Questions this paper answers
- If you delete a chatbot's own earlier replies from the conversation history, do its later answers get worse?
- How does assistant-turn omission in the dialogue history affect judged response quality and on-topic rate across instruction-tuned LLMs?
- How do I trim assistant turns out of a multi-turn prompt without degrading the next response?
- Should I keep past model replies in my chat history, or send only the user turns?
- Omitting all prior assistant responses preserves average response quality for DeepSeek-R1-Distill-Llama-8B and GPT-OSS-20B, while Qwen3-4B and GPT-5.2 lose some quality relative to full context. Responses stay on-topic under omission for all four models.
Holds for: 350 English multi-turn conversations, 200 from WildChat and 150 from ShareLM, 5-10 rounds each, spanning creative-writing, science and coding keyword categories; GPT-5 as pairwise judge seeing prior user and assistant turns.
- When the GPT-5 judge is shown only prior user-side turns, assistant-omitted responses often perform comparably to full-context responses on both the quality and on-topic dimensions across the four models.
Holds for: Same 350 WildChat and ShareLM conversations and same four models; a judge without assistant history cannot resolve prompts that explicitly reference an earlier response.
- How much of the prompt do you save in a long chat by trimming the chatbot's own earlier replies?
- What context-length reduction do the reduced-context configurations yield over full dialogue history as a conversation grows?
- How do I cut token spend on a growing chat history without truncating or summarizing?
- Will dropping assistant turns actually make my multi-turn requests meaningfully cheaper?
- Cumulative context length grows linearly with conversation length under full context. The summarized, last-turn-only and assistant-omitted configurations stay relatively constant, so a reduced configuration that matches full-context quality uses roughly 8x less context.
Holds for: Cumulative context measured in characters over 5-10 round WildChat and ShareLM conversations, plotted from GPT-5.2 generations; does not include generated response length.
- In real conversations with a chatbot, how often is a person's next message a brand-new request rather than a follow-up?
- What proportion of non-initial user turns in in-the-wild human-LLM logs are self-contained versus dependent on prior assistant output?
- How do I tell whether a multi-turn chat corpus really tests long-context multi-turn reasoning?
- Can I trust WildChat-style logs as a benchmark for multi-turn context dependence?
- In real-world multi-turn chats, 36.4% of non-initial user turns are self-contained new asks, 30.5% are follow-ups with concrete feedback, and 33.1% are follow-ups referencing an earlier turn without actionable feedback.
Holds for: GPT-5 as automated annotator over the sampled WildChat and ShareLM English conversations of 5-10 rounds; the 33.1% is an upper bound on assistant-dependence.
- "Do LLMs Benefit From Their Own Words?" argues that in-the-wild multi-turn chat logs are weak benchmarks for long-context multi-turn reasoning, and calls for corpora curated for genuine multi-turn dependence. The reason is that a large share of turns in real chat logs do not depend on earlier assistant responses.
Holds for: Argued from English conversations of 5-10 rounds sampled from WildChat and ShareLM across creative-writing, science, coding and other keyword categories; agentic settings with tool outputs and scratchpads are discussed but not measured.
- Which kinds of user messages still need the chatbot's earlier answer in the history, and which do not?
- Does the benefit of retaining assistant history differ between new-ask turns and follow-up turns?
- How do I decide per turn whether a user message needs the assistant's prior response?
- If most of my users ask follow-up questions, is stripping assistant history still safe for me?
- For Qwen3-4B and GPT-5.2, assistant-side history is most beneficial for follow-up turns, while full-context and assistant-omitted prompting perform comparably on new-ask turns.
Holds for: The two models for which uniform omission lowered overall quality; win rates averaged over the quality and on-topic dimensions, with stars marking significant differences.
- In real-world multi-turn chats, 36.4% of non-initial user turns are self-contained new asks, 30.5% are follow-ups with concrete feedback, and 33.1% are follow-ups referencing an earlier turn without actionable feedback.
Holds for: GPT-5 as automated annotator over the sampled WildChat and ShareLM English conversations of 5-10 rounds; the 33.1% is an upper bound on assistant-dependence.
- Can a chatbot's own earlier mistakes in the conversation get repeated into later answers?
- What does context pollution from retained assistant turns look like in multi-turn LLM conversations?
- How do I stop a wrong assumption or hallucination from an earlier reply carrying into later turns?
- Is keeping my model's previous outputs in the prompt risking repeated hallucinations and stale code?
- Retained assistant responses produce context pollution, with UMAP arguments carried into t-SNE code, a hallucinated fact about a novel repeated across turns, and a misattributed NBER citation. An earlier response's style also overrode a reflection request, and a temperature formula was reused incorrectly.
Holds for: Select cases the authors reviewed among turns an LLM annotator flagged as polluted; illustrative rather than a frequency estimate, and includes GPT-5.2 as well as smaller open models.
- Can a small model learn when to keep the chatbot's earlier replies and when to throw them away?
- Can a per-turn classifier select between full and assistant-omitted context, and is judge preference predictable from turn features?
- How do I build a per-turn policy that drops assistant history only when it is safe?
- Is adaptive per-turn context filtering worth implementing, or will I lose quality against always sending full history?
- A per-turn classifier choosing between full and assistant-omitted context retains over 99% of full-context-only performance while using an average of 87% of the total token cost. It beats a heuristic that omits assistant history on all new-ask turns.
Holds for: GPT-5.2 with its own responses fed back into context; L1-regularized logistic regression over round number, context lengths, prompt category and PCA-reduced embeddings; the heuristic baseline uses a 20% held-out subset.
- Predicting whether the judge will prefer full context over assistant-omitted context is hard: the L1-regularized logistic regression reaches only a 5-fold cross-validated F1 of 0.6106 ± 0.0119. None of the top 20 features are significant at the 5% level.
Holds for: GPT-5.2 responses on the sampled WildChat and ShareLM conversations; features are round number, context lengths, prompt category, and 20 principal components each of the prompt and history embeddings.
- Is it better to replace a chatbot's past answers with one-sentence summaries than to keep them in full?
- How does one-sentence summarization of assistant turns compare with full history and with full omission across four models on WildChat and ShareLM?
- How do I compress assistant turns in a chat history instead of deleting them outright?
- Should I summarize my model's earlier replies rather than dropping them or keeping them verbatim?
- Replacing each assistant response with a one-sentence summary is the most effective of the reduced-context configurations, often exceeding full context across the four models. It also cuts median response length by roughly 25%.
Holds for: Four models on the 350 WildChat and ShareLM conversations; the summary replaces each prior assistant turn in place, and the median-length reduction is measured over generated responses.
- How much can you trust a big model's scoring of which chat answer is better, and does what it sees change its verdict?
- How well does the GPT-5 judge agree with manual annotation, and does restricting the judge to user-side turns change the preference between context conditions?
- How do I set up an LLM judge for comparing responses under different context conditions without biasing it toward the longer history?
- Can I rely on LLM-as-judge results comparing full-context and assistant-omitted responses?
- The GPT-5 judge agreed with an author's manual verdict on 74 of 80 quality judgments (92.5%) and 75 of 80 on-topic judgments (93.75%).
Holds for: 20 sampled judgments per model across the 4 models, annotated by one author; no multi-annotator human study was run.
- When the GPT-5 judge is shown only prior user-side turns, assistant-omitted responses often perform comparably to full-context responses on both the quality and on-topic dimensions across the four models.
Holds for: Same 350 WildChat and ShareLM conversations and same four models; a judge without assistant history cannot resolve prompts that explicitly reference an earlier response.
- Is there a study that asks whether keeping a chatbot's own past replies in the prompt is actually worth it?
- What work evaluates assistant-history retention on in-the-wild human-LLM logs rather than synthetic multi-turn dialogues?
- What should I read before designing context management for a multi-turn chat or agent system?
- Which paper should I cite if I want to argue that real chat logs are weak multi-turn benchmarks?
- Huang et al.'s "Do LLMs Benefit From Their Own Words?" questions the default assumption of multi-turn chat and agentic context management that retaining a model's own past responses reliably helps. It tests that assumption on in-the-wild human-LLM logs rather than synthetic dialogues.
Holds for: As of the August 2026 version of the preprint; prior turn-level context-editing work such as ERGO evaluated on synthetic conversations, and earlier conversational-QA findings about irrelevant turns concerned human-human histories.
- "Do LLMs Benefit From Their Own Words?" argues that in-the-wild multi-turn chat logs are weak benchmarks for long-context multi-turn reasoning, and calls for corpora curated for genuine multi-turn dependence. The reason is that a large share of turns in real chat logs do not depend on earlier assistant responses.
Holds for: Argued from English conversations of 5-10 rounds sampled from WildChat and ShareLM across creative-writing, science, coding and other keyword categories; agentic settings with tool outputs and scratchpads are discussed but not measured.
- Do big commercial chatbots get thrown off by their own earlier replies, or is that only a small-model problem?
- Which models were evaluated for assistant-turn omission, and does over-conditioning on prior assistant output appear in frontier models too?
- How do I check whether the model I use is being misled by its own earlier turns?
- I run a frontier model in production, is over-conditioning on its own past answers something I have to worry about?
- Omitting all prior assistant responses preserves average response quality for DeepSeek-R1-Distill-Llama-8B and GPT-OSS-20B, while Qwen3-4B and GPT-5.2 lose some quality relative to full context. Responses stay on-topic under omission for all four models.
Holds for: 350 English multi-turn conversations, 200 from WildChat and 150 from ShareLM, 5-10 rounds each, spanning creative-writing, science and coding keyword categories; GPT-5 as pairwise judge seeing prior user and assistant turns.
- Retained assistant responses produce context pollution, with UMAP arguments carried into t-SNE code, a hallucinated fact about a novel repeated across turns, and a misattributed NBER citation. An earlier response's style also overrode a reflection request, and a temperature formula was reused incorrectly.
Holds for: Select cases the authors reviewed among turns an LLM annotator flagged as polluted; illustrative rather than a frequency estimate, and includes GPT-5.2 as well as smaller open models.
Claims and scope
- Omitting all prior assistant responses preserves average response quality for DeepSeek-R1-Distill-Llama-8B and GPT-OSS-20B, while Qwen3-4B and GPT-5.2 lose some quality relative to full context. Responses stay on-topic under omission for all four models. (Figure 2 (third row))
Scope: 350 English multi-turn conversations, 200 from WildChat and 150 from ShareLM, 5-10 rounds each, spanning creative-writing, science and coding keyword categories; GPT-5 as pairwise judge seeing prior user and assistant turns.
- When the GPT-5 judge is shown only prior user-side turns, assistant-omitted responses often perform comparably to full-context responses on both the quality and on-topic dimensions across the four models. (Figure 10 (Section A.11))
Scope: Same 350 WildChat and ShareLM conversations and same four models; a judge without assistant history cannot resolve prompts that explicitly reference an earlier response.
- Cumulative context length grows linearly with conversation length under full context. The summarized, last-turn-only and assistant-omitted configurations stay relatively constant, so a reduced configuration that matches full-context quality uses roughly 8x less context. (Figure 4 (center))
Scope: Cumulative context measured in characters over 5-10 round WildChat and ShareLM conversations, plotted from GPT-5.2 generations; does not include generated response length.
- In real-world multi-turn chats, 36.4% of non-initial user turns are self-contained new asks, 30.5% are follow-ups with concrete feedback, and 33.1% are follow-ups referencing an earlier turn without actionable feedback. (Section A.7)
Scope: GPT-5 as automated annotator over the sampled WildChat and ShareLM English conversations of 5-10 rounds; the 33.1% is an upper bound on assistant-dependence.
- For Qwen3-4B and GPT-5.2, assistant-side history is most beneficial for follow-up turns, while full-context and assistant-omitted prompting perform comparably on new-ask turns. (Figure 8 (Section A.7))
Scope: The two models for which uniform omission lowered overall quality; win rates averaged over the quality and on-topic dimensions, with stars marking significant differences.
- Retained assistant responses produce context pollution, with UMAP arguments carried into t-SNE code, a hallucinated fact about a novel repeated across turns, and a misattributed NBER citation. An earlier response's style also overrode a reflection request, and a temperature formula was reused incorrectly. (Table 1 (examples in Section A.22))
Scope: Select cases the authors reviewed among turns an LLM annotator flagged as polluted; illustrative rather than a frequency estimate, and includes GPT-5.2 as well as smaller open models.
- A per-turn classifier choosing between full and assistant-omitted context retains over 99% of full-context-only performance while using an average of 87% of the total token cost. It beats a heuristic that omits assistant history on all new-ask turns. (Figure 6 (Section 5))
Scope: GPT-5.2 with its own responses fed back into context; L1-regularized logistic regression over round number, context lengths, prompt category and PCA-reduced embeddings; the heuristic baseline uses a 20% held-out subset.
- Predicting whether the judge will prefer full context over assistant-omitted context is hard: the L1-regularized logistic regression reaches only a 5-fold cross-validated F1 of 0.6106 ± 0.0119. None of the top 20 features are significant at the 5% level. (Table 5 (Section A.15))
Scope: GPT-5.2 responses on the sampled WildChat and ShareLM conversations; features are round number, context lengths, prompt category, and 20 principal components each of the prompt and history embeddings.
- Replacing each assistant response with a one-sentence summary is the most effective of the reduced-context configurations, often exceeding full context across the four models. It also cuts median response length by roughly 25%. (Figure 2 (top row), Section 3)
Scope: Four models on the 350 WildChat and ShareLM conversations; the summary replaces each prior assistant turn in place, and the median-length reduction is measured over generated responses.
- The GPT-5 judge agreed with an author's manual verdict on 74 of 80 quality judgments (92.5%) and 75 of 80 on-topic judgments (93.75%). (Section A.5)
Scope: 20 sampled judgments per model across the 4 models, annotated by one author; no multi-annotator human study was run.
- Huang et al.'s "Do LLMs Benefit From Their Own Words?" questions the default assumption of multi-turn chat and agentic context management that retaining a model's own past responses reliably helps. It tests that assumption on in-the-wild human-LLM logs rather than synthetic dialogues.
Scope: As of the August 2026 version of the preprint; prior turn-level context-editing work such as ERGO evaluated on synthetic conversations, and earlier conversational-QA findings about irrelevant turns concerned human-human histories.
- "Do LLMs Benefit From Their Own Words?" argues that in-the-wild multi-turn chat logs are weak benchmarks for long-context multi-turn reasoning, and calls for corpora curated for genuine multi-turn dependence. The reason is that a large share of turns in real chat logs do not depend on earlier assistant responses.
Scope: Argued from English conversations of 5-10 rounds sampled from WildChat and ShareLM across creative-writing, science, coding and other keyword categories; agentic settings with tool outputs and scratchpads are discussed but not measured.
Common misreadings
- Omitting assistant history is not shown to be universally free: for Qwen3-4B and GPT-5.2, uniformly dropping assistant responses lowers average judged quality under a judge that sees the full history.
- The 36.4% figure counts non-initial user turns classified as self-contained new asks by a GPT-5 annotator on filtered English conversations from WildChat and ShareLM; it is not a claim about all chat traffic.
- The context-pollution cases are select examples the authors reviewed among annotator-flagged turns, so they establish that the failure mode exists and reaches GPT-5.2. The paper measures how often it occurs separately, in Table 1.
- The adaptive omission strategy is not a strong predictive model: its 5-fold cross-validated F1 is 0.6106 and no individual feature is significant, so its context savings come from a weak signal rather than a reliable per-turn prediction.
- The roughly 8x reduction is in cumulative context characters over conversation rounds. No speedup or dollar cost saving was measured.
Terminology in this paper
- Assistant-Omitted (AO) context
- A prompting configuration in which every past assistant response in a multi-turn conversation is replaced by the placeholder "[Response provided]", so the model conditions only on prior user turns while the alternating user/assistant structure is preserved.
- context pollution
- The phenomenon in which a language model over-conditions on its own earlier responses, so errors, hallucinations or stylistic artifacts introduced in one turn propagate into later turns.
- New Ask
- A non-initial user turn that introduces a fully self-contained request, understandable without any prior conversation round.
- Follow-up with Feedback
- A user turn that gives concrete, actionable feedback on a prior assistant response, such as "use Python instead of Java for the code example".
- Follow-up without Feedback
- A user turn that refers to an earlier conversation round without any concrete instruction for revision, such as "reflect on your response".
How to cite
@misc{huang2026llmsbenefitwords,
title={Do LLMs Benefit From Their Own Words?},
author={Jenny Y. Huang and Leshem Choshen and Ramon Astudillo and Tamara Broderick and Jacob Andreas},
year={2026},
eprint={2602.24287},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2602.24287},
}
References
See the full reference list in the paper.