检测矛盾不等于安全解决,大模型在多轮推理中仍会忽略证据冲突。
Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs

- 通过多轮文档累积测试,发现模型能识别矛盾却无法据此修正输出。
- 超过5万次评估显示,单轮检测结果严重高估了RAG系统的安全性。
- 问题根源在行为决策环节,关键信息虽被关注却未影响最终回答。
检索增强型大模型用于依赖证据质量的高风险任务,但现有评估假设单轮鲁棒性可推及多轮累积证据场景。我们证明该假设根本错误:模型存在监测-控制鸿沟——能识别矛盾证据,却无法据此约束最终推荐。在四个模型家族(1.5B-32B参数)上,通过超过5万次轮次级评估验证:单轮诊断系统性高估RAG安全性;矛盾识别与安全解决无相关性;人工验证支持此现象。无通用提示修复方案。多源机制分析(隐状态探测、注意力分析、响应策略分类)指向行为选择为缺陷核心:危险相关信息在生成时被内部表征并获得更强注意力,却未能抑制不当输出。在高风险场景中部署前,必须测量并弥合这一认知与行为间的差距。
原文摘要 · Abstract (English)
Retrieval-augmented LLMs are deployed for tasks where evidence quality determines action safety, yet evaluation protocols assume that single-turn robustness predicts robustness when evidence accumulates across turns. We show this assumption is fundamentally incorrect. Models exhibit a monitoring-control gap: they readily acknowledge contradictory evidence, yet this awareness fails to constrain their final recommendations - detecting epistemic conflict does not imply resolving it safely. Through a multi-turn document accumulation protocol across four model families (1.5B-32B parameters) and over 50,000 turn-level evaluations, we demonstrate that single-turn diagnostics systematically overestimate RAG safety, that contradiction acknowledgement is uncorrelated with safe resolution, a pattern corroborated by targeted human validation, and that no universal prompt fix exists. Converging mechanism evidence - hidden-state probing, attention analysis, and response-strategy taxonomy - points to action selection as the most plausible locus of the deficit: danger-relevant information is internally represented and receives enhanced attention during unsafe generation, yet fails to constrain output behavior. The gap between what models recognize and what they do must be measured and closed before retrieval-augmented systems can be trusted in high-stakes settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。