提升长文本事实性评估,精准捕捉缺失和关系型事实。
VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Facts
- 通过识别与补全不完整事实,改进事实提取精度。
- 在FactRBench上验证,大模型同时提升精确率与召回率。
- 适合需要全面评估生成内容真实性的研究者使用。
大型语言模型(LLMs)在生成长文本方面表现优异,但其事实性评估因生成内容中句子间的复杂依赖关系而困难。现有方法多采用分解-去上下文-验证的流程,常忽略关键上下文信息并遗漏重要关系事实。本文提出VeriFact框架,通过识别和修复不完整及缺失的事实,提升事实提取质量,从而支持更准确的事实验证。同时引入FactRBench基准,评估长文本生成中的精确率与召回率,此前研究主要关注精确率。FactRBench提供来自先进LLMs和人工撰写的参考事实集,支持召回率评估。实证结果表明,VeriFact显著提升了事实完整性,并保留了包含关键关系信息的复杂事实,使事实性评估更准确。在FactRBench上对多种开源与闭源模型的评测显示,同模型族中更大规模模型在精确率和召回率上均提升,但高精确率未必对应高召回率,凸显全面评估的重要性。
原文摘要 · Abstract (English)
Large language models (LLMs) excel at generating long-form responses, but evaluating their factuality remains challenging due to complex inter-sentence dependencies within the generated facts. Prior solutions predominantly follow a decompose-decontextualize-verify pipeline but often fail to capture essential context and miss key relational facts. In this paper, we introduce VeriFact, a factuality evaluation framework designed to enhance fact extraction by identifying and resolving incomplete and missing facts to support more accurate verification results. Moreover, we introduce FactRBench , a benchmark that evaluates both precision and recall in long-form model responses, whereas prior work primarily focuses on precision. FactRBench provides reference fact sets from advanced LLMs and human-written answers, enabling recall assessment. Empirical evaluations show that VeriFact significantly enhances fact completeness and preserves complex facts with critical relational information, resulting in more accurate factuality evaluation. Benchmarking various open- and close-weight LLMs on FactRBench indicate that larger models within same model family improve precision and recall, but high precision does not always correlate with high recall, underscoring the importance of comprehensive factuality assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。