arXiv:2510.07083cs.CL2025-10被引 6

为关键信息错误设计更敏感的模型事实性评估方法

All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations

  • 基于查询重要性加权判断每条陈述的真实度
  • 在6733个测试用例中显著提升关键错误检测率
  • 适合关注大模型可靠性与评估体系改进的研究者

现有大语言模型(LLM)事实性评估方法将所有陈述视为同等重要,导致关键信息缺失或错误时评估结果失真。为此,我们构建了VITALERRORS数据集,包含6,733个经微调的查询与响应,专门模拟关键信息的遗漏或篡改。实验表明,现有评估指标对关键信息错误不敏感。为此提出VITAL评估体系,通过引入陈述与查询的相关性与重要性权重,增强对关键错误的识别能力。分析显示,VITAL在检测关键信息错误方面优于以往方法。该数据集、度量标准与分析为更准确可靠的LLM事实性评估提供了基础。

原文摘要 · Abstract (English)

Existing methods for evaluating the factuality of large language model (LLM) responses treat all claims as equally important. This results in misleading evaluations when vital information is missing or incorrect as it receives the same weight as peripheral details, raising the question: how can we reliably detect such differences when there are errors in key information? Current approaches that measure factuality tend to be insensitive to omitted or false key information. To investigate this lack of sensitivity, we construct VITALERRORS, a benchmark of 6,733 queries with minimally altered LLM responses designed to omit or falsify key information. Using this dataset, we demonstrate the insensitivities of existing evaluation metrics to key information errors. To address this gap, we introduce VITAL, a set of metrics that provide greater sensitivity in measuring the factuality of responses by incorporating the relevance and importance of claims with respect to the query. Our analysis demonstrates that VITAL metrics more reliably detect errors in key information than previous methods. Our dataset, metrics, and analysis provide a foundation for more accurate and robust assessment of LLM factuality.

事实性评估大模型评测重要性加权基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。