评测大模型在知识冲突下的决策能力,发现无一模型能稳定应对各类矛盾。
KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

- 构建238个动态交互任务,模拟知识冲突、输入不一致与时间矛盾场景
- 九个模型在事实修正、身份一致与时间冲突上均表现不稳定,最高准确率不足60%
- 适合研究模型推理鲁棒性、安全防护机制的开发者和研究人员
随着大模型越来越多地通过工具执行任务,必须在用户指令、参数化知识和动态环境观测之间进行协调后才能行动。我们提出KC-Bench,一个受控的多轮基准测试,用于衡量模型在世界知识冲突、输入不一致和多源时间冲突下的处理能力。该基准包含238个经过人工筛选的任务,源自超过1000个生成候选任务,融合用户模拟器、状态化工具、确定性环境断言、开源自然语言评估器及人工轨迹验证。对九个模型(包括DeepSeek-V4-Flash、GLM-5.2、MiniMax-M3)的评估显示,跨领域表现差异显著:无一模型能在所有场景下可靠完成事实修正、身份一致性检查和时间冲突解决。在模拟环境中,未察觉的冲突可能传播至工具调用或合成敏感数据流。KC-Bench聚焦于模型层面的行为诊断,而非完整智能体框架的排名,为开发具备冲突感知的推理与执行防护机制提供可复现的分析工具。
原文摘要 · Abstract (English)
As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts. Its 238 tasks are manually screened from more than 1,000 generated candidates and combine a user simulator, stateful tools, deterministic environment assertions, an open-source natural-language evaluator, and human trajectory verification. Evaluation of nine models, including DeepSeek-V4-Flash, GLM-5.2, and MiniMax-M3, shows substantial cross-domain variation: no model handles factual correction, identity consistency checking, and temporal conflict resolution reliably across all settings. In the simulated environments, missed conflicts can propagate to tool calls or synthetic protected-data flows. KC-Bench isolates this model-level behavior rather than ranking complete agent frameworks, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。