提出兼顾准确性和完整性的长文本事实评估新方法。
Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation
- 联合评估生成内容的准确率与覆盖度,突破仅看准确的局限。
- 发现当前大模型在事实覆盖上远低于准确率,存在明显遗漏。
- 引入重要性加权,更关注关键事实是否被覆盖,适合质量评估者使用。
评估大语言模型生成的长文本事实准确性仍具挑战性,尤其当回应开放且包含大量细粒度事实时。现有方法主要关注准确率:将回答分解为原子陈述,并与维基百科等外部知识源验证每项陈述。但忽略了同样重要的维度——召回率,即生成内容是否涵盖应包含的相关事实。本文提出一个综合的事实评估框架,同时衡量准确率与召回率。该方法利用外部知识源构建参考事实,并判断其是否被生成文本覆盖。进一步引入基于相关性和显著性的加权机制。分析表明,当前大模型在准确率上表现远优于召回率,说明事实不完整性仍是长文本生成的主要瓶颈,模型更擅长覆盖高度重要的事实而非全部相关事实。
原文摘要 · Abstract (English)
Evaluating the factuality of long-form output generated by large language models (LLMs) remains challenging, particularly when responses are open-ended and contain many fine-grained factual statements. Existing evaluation methods primarily focus on precision: they decompose a response into atomic claims and verify each claim against external knowledge sources such as Wikipedia. However, this overlooks an equally important dimension of factuality: recall, whether the generated response covers the relevant facts that should be included. We propose a comprehensive factuality evaluation framework that jointly measures precision and recall. Our method leverages external knowledge sources to construct reference facts and determine whether they are captured in generated text. We further introduce an importance-aware weighting scheme based on relevance and salience. Our analysis reveals that current LLMs perform substantially better on precision than on recall, suggesting that factual incompleteness remains a major limitation of long-form generation and that models are generally better at covering highly important facts than the full set of relevant facts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。