提出新指标评估大模型回答是否真正依赖外部上下文。
ConSens: Assessing context grounding in open-book question answering
- 通过对比有无上下文时的困惑度差异衡量模型依赖程度。
- 实验验证该方法能有效识别答案是否基于提供上下文。
- 计算高效、可解释,适合各类开放问答系统评估。
大型语言模型在开放问答任务中表现出色,但关键挑战在于确保模型回答基于提供的外部上下文,而非其内部参数知识(可能过时或错误)。现有基于LLM作为裁判的评估方法存在偏见、可扩展性差及依赖昂贵外部系统等局限。为此,本文提出一种新指标:对比模型在有无上下文条件下的回答困惑度差异,量化其对上下文的依赖程度。实验表明该指标能有效判断答案是否真正基于给定上下文。相比现有方法,该指标计算高效、可解释性强,且适用于多种场景,为开放问答系统中的上下文利用评估提供了可扩展、实用的解决方案。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated considerable success in open-book question answering (QA), where the task requires generating answers grounded in a provided external context. A critical challenge in open-book QA is to ensure that model responses are based on the provided context rather than its parametric knowledge, which can be outdated, incomplete, or incorrect. Existing evaluation methods, primarily based on the LLM-as-a-judge approach, face significant limitations, including biases, scalability issues, and dependence on costly external systems. To address these challenges, we propose a novel metric that contrasts the perplexity of the model response under two conditions: when the context is provided and when it is not. The resulting score quantifies the extent to which the model's answer relies on the provided context. The validity of this metric is demonstrated through a series of experiments that show its effectiveness in identifying whether a given answer is grounded in the provided context. Unlike existing approaches, this metric is computationally efficient, interpretable, and adaptable to various use cases, offering a scalable and practical solution to assess context utilization in open-book QA systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。