构建多跳图结构基准,测试大模型在矛盾信息中的推理能力
MAGIC: A Multi-Hop and Graph-Based Benchmark for Inter-Context Conflicts in Retrieval-Augmented Generation
- 基于知识图谱生成多样且细微的上下文冲突
- 大模型在需要多跳推理时检测冲突准确率低
- 适合研究大模型信息整合与错误溯源的学者
检索增强生成(RAG)系统中常出现知识冲突,即检索文档之间不一致或与模型参数化知识相矛盾。现有基准存在局限:聚焦问答任务、依赖实体替换、冲突类型单一。为此,我们提出基于知识图谱(KG)的框架,通过显式关系结构生成多样且微妙的冲突,同时保证可解释性。在新基准MAGIC上的实验揭示了大模型处理知识冲突的内在机制:开源与商用模型在多跳推理场景下均难以有效检测冲突,且常无法准确定位矛盾来源。深入分析为提升大模型整合异构甚至矛盾信息的能力提供了基础。
原文摘要 · Abstract (English)
Knowledge conflict often arises in retrieval-augmented generation (RAG) systems, where retrieved documents may be inconsistent with one another or contradict the model's parametric knowledge. Existing benchmarks for investigating the phenomenon have notable limitations, including a narrow focus on the question answering setup, heavy reliance on entity substitution techniques, and a restricted range of conflict types. To address these issues, we propose a knowledge graph (KG)-based framework that generates varied and subtle conflicts between two similar yet distinct contexts, while ensuring interpretability through the explicit relational structure of KGs. Experimental results on our benchmark, MAGIC, provide intriguing insights into the inner workings of LLMs regarding knowledge conflict: both open-source and proprietary models struggle with conflict detection -- especially when multi-hop reasoning is required -- and often fail to pinpoint the exact source of contradictions. Finally, we present in-depth analyses that serve as a foundation for improving LLMs in integrating diverse, sometimes even conflicting, information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。