让大模型学会对检索信息保持怀疑,提升高风险场景可靠性
Can LLMs Take Retrieved Information with a Grain of Salt?

- 通过提醒先验知识、校准确定性、简化上下文提升模型对不确定性的判断力
- 平均降低25%的响应偏差,无需修改模型参数
- 适合医疗金融等需谨慎决策的领域使用
大语言模型在检索增强任务中表现优异,但在如何根据检索信息的确定性调整回应方面仍存在系统性不足,这在医疗、金融等高风险领域可能带来严重后果。我们评估了八种LLM在上下文确定性服从性上的表现,发现其在观察不确定性上下文后难以回忆先验知识、误读确定性表达、过度信任复杂上下文。为此,我们提出一种交互策略:结合先验提醒、确定性重校和上下文简化。该策略平均减少25%的服从误差,且不需修改模型权重,验证了交互设计对提升LLM可靠性的有效性。贡献包括一个可量化的评估指标、关于LLM不确定性处理的实证洞察,以及一种可跨模型复用的改进策略。
原文摘要 · Abstract (English)
Large language models have demonstrated impressive retrieval-augmented capabilities. However, a crucial area remains underexplored: their ability to appropriately adapt responses to the certainty of the retrieved information. It is a limitation with real consequences in high-stakes domains like medicine and finance. We evaluate eight LLMs on their context-certainty obedience, measuring how well they adjust responses to match expressed context certainty. Our analysis reveals systematic limitations: LLMs struggle to recall prior knowledge after observing an uncertain context, misinterpret expressed certainties, and overtrust complex contexts. To address these, we propose an interaction strategy combining prior reminders, certainty recalibration, and context simplification. This approach reduces obedience errors by 25% on average, without modifying model weights, demonstrating the efficacy of interaction design in enhancing LLM reliability. Our contributions include a principled evaluation metric, empirical insights into LLMs' uncertainty handling, and a portable strategy to improve context-certainty obedience across diverse LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。