提出新评估方法,精准衡量模型用证据答对题的能力。
Diagnosing Evidence Utilization in Long-Context and Retrieval-Augmented Language Models under Matched Evidence Conditions
- 设计四条件对照实验,分离模型真实利用证据的能力
- 发现长上下文输入比检索增强更易丢失关键信息
- 适合关注大模型推理可信度的研究者和开发者
最终答案准确率、检索召回率与引用重合度无法揭示长上下文或检索增强语言模型从给定证据中实际获得的答案优势。模型可能依赖参数先验、未能使用存在证据,或引用相关文本但未转化为答案。本文提出一种四条件诊断协议,在匹配的样本、模型、提示和评分规则下评估证据利用情况。该协议对比无证据、全上下文、检索证据与最优参考(oracle)四种条件,采用归一化最优参考证据利用率(ONCU)作为有效基准。实证研究在18,000个符合ONCU标准的预测上测试了来自Qwen、Gemma、Llama和Mistral系列的五个本地开源模型,覆盖Controlled-ONCU-safe16K、HotpotQA-ONCU和2WikiMultiHopQA-ONCU数据集。结果表明任务依赖性明显:合成环境中,相同证据嵌入长输入时恢复能力下降;真实多跳任务中,全上下文输入在无需归一化的答案与证据指标上优于检索输入,且ONCU支持相同趋势。强化检索设置虽缩小部分差距,但未改变整体结论。核心贡献并非单一利用率数值,而是一套匹配诊断协议,可区分无证据回答能力、最优证据可恢复性、全上下文恢复、检索条件恢复、分母有效性及配套答案/证据诊断。
原文摘要 · Abstract (English)
Final-answer accuracy, retrieval recall, and citation overlap do not reveal how much answer advantage a long-context or retrieval-augmented language model actually recovers from supplied evidence. A model may answer from parametric priors, fail to use evidence that is present, or cite relevant text without converting it into the final answer. This paper introduces a four-condition diagnostic protocol for evidence-utilization evaluation under matched examples, models, prompts, and scoring rules. The protocol compares no-evidence, full-context, retrieved-evidence, and oracle-evidence reference conditions, and uses Oracle-Reference Normalized Context Utilization (ONCU) as a denominator-valid estimate of recovered oracle-reference evidence advantage. The empirical study evaluates five local open-weight models from the Qwen, Gemma, Llama, and Mistral families over Controlled-ONCU-safe16K, HotpotQA-ONCU, and 2WikiMultiHopQA-ONCU, comprising 18,000 ONCU-compatible predictions. Results show a task-dependent diagnostic pattern: controlled synthetic settings expose reduced recovery when the same evidence is embedded in long input rather than supplied compactly, while realistic multi-hop reconstructions show that full-context inputs outperform the tested retrieved inputs in denominator-free answer and evidence metrics, with ONCU supporting the same direction on oracle-improving groups. Sensitivity audits with stronger retrieval settings narrow some gaps but do not overturn the scoped interpretation. The main contribution is therefore not a single utilization ratio, but a matched diagnostic protocol that separates no-evidence answerability, oracle-evidence recoverability, full-context recovery, retrieval-conditioned recovery, denominator validity, and companion answer/evidence diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。