发现大模型易受少量不当内容干扰,提出新方法提升其在复杂上下文中的可靠性。
Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts
- 用神经科学模型模拟上下文竞争,发现模型更倾向采纳少见信息
- 实测显示少量不当内容可使输出质量下降,且模型行为呈负面趋势
- 新方法无需大量标注数据,能有效识别并过滤不当信号
引入外部上下文可显著提升大语言模型(LLM)的响应质量,但现实场景中常混杂相关与不恰当内容,带来可靠性风险。为研究此问题,我们构建了包含真实世界混合上下文的「中毒上下文测试床」,并借鉴动物联想学习理论,将神经科学中的Rescorla-Wagner(RW)模型引入量化分析。结果揭示一致行为模式:模型倾向于采纳上下文中较少出现的信息。这一倾向在真实环境中有害,即微小比例的不恰当内容即可显著降低输出质量。实证评估验证了该脆弱性。为此,我们提出基于两阶段微调的RW-Steering方法,使模型内生识别并忽略不恰当信号。相比依赖大量标注数据的已有方法,本方法在不同比例的不恰当内容下仍具强泛化能力。实验表明,最优微调模型使响应质量提升39.8%,逆转不良行为曲线,确立其作为提升实际应用中大模型可信度的稳健通用解决方案。
原文摘要 · Abstract (English)
Incorporating external context can significantly enhance the response quality of Large Language Models (LLMs). However, real-world contexts often mix relevant information with disproportionate inappropriate content, posing reliability risks. How do LLMs process and prioritize mixed context? To study this, we introduce the Poisoned Context Testbed, pairing queries with real-world contexts containing relevant and inappropriate content. Inspired by associative learning in animals, we adapt the Rescorla-Wagner (RW) model from neuroscience to quantify how competing contextual signals influence LLM outputs. Our adapted model reveals a consistent behavioral pattern: LLMs exhibit a strong tendency to incorporate information that is less prevalent in the context. This susceptibility is harmful in real-world settings, where small amounts of inappropriate content can substantially degrade response quality. Empirical evaluations on our testbed further confirm this vulnerability. To tackle this, we introduce RW-Steering, a two-stage finetuning-based approach that enables the model to internally identify and ignore inappropriate signals. Unlike prior methods that rely on extensive supervision across diverse context mixtures, RW-Steering generalizes robustly across varying proportions of inappropriate content. Experiments show that our best fine-tuned model improves response quality by 39.8% and reverses the undesirable behavior curve, establishing RW-Steering as a robust, generalizable context engineering solution for improving LLM safety in real-world use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。