arXiv:2509.16534cs.CLcs.AI2025-09EMNLP被引 3

提出整合式接地评估框架,揭示大模型在多证据验证中的弱点与改进方向。

InteGround: On the Evaluation of Verification and Retrieval Planning in Integrative Grounding

  • 构建多源证据整合的验证与检索规划评估体系
  • 发现模型在信息不全时会依赖内部知识进行合理化解释
  • 自省能力可稳定提升接地质量,前提归纳优于盲目检索

将大语言模型(LLMs)与外部知识源结合是实现可信预测的有前景方法。现有方法对简单查询表现良好,但现实信息需求常需综合多个相互依赖的证据。本文提出“整合式接地”——即为支持假设性问题,检索并验证多条关联证据的挑战。为系统研究此问题,我们复用四个领域数据构建评估基准。研究发现:在可信赖性验证中,尽管模型对冗余证据具有鲁棒性,但在信息缺失时倾向于使用内部知识进行合理化;在检索规划策略方面,无目的性规划会引入噪声导致性能下降,而基于前提归纳的方法因具备逻辑约束表现出色。此外,模型零样本自省能力始终能提升接地质量。这些发现为构建更有效的整合式接地系统提供了关键指导。

原文摘要 · Abstract (English)

Grounding large language models (LLMs) in external knowledge sources is a promising method for faithful prediction. While existing grounding approaches work well for simple queries, many real-world information needs require synthesizing multiple pieces of evidence. We introduce "integrative grounding" -- the challenge of retrieving and verifying multiple inter-dependent pieces of evidence to support a hypothesis query. To systematically study this problem, we repurpose data from four domains for evaluating integrative grounding capabilities. Our investigation reveals two critical findings: First, in groundedness verification, while LLMs are robust to redundant evidence, they tend to rationalize using internal knowledge when information is incomplete. Second, in examining retrieval planning strategies, we find that undirected planning can degrade performance through noise introduction, while premise abduction emerges as a promising approach due to its logical constraints. Additionally, LLMs' zero-shot self-reflection capabilities consistently improve grounding quality. These insights provide valuable direction for developing more effective integrative grounding systems.

大模型知识融合推理验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。