评估长文本生成中事实覆盖多样性,比单纯查对真假更全面。
Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation
- 将长文本拆解为原子事实,逐条验证并匹配预期主题维度。
- 在TREC和ClueWeb数据集上与人工判断高度相关,可衡量多方面覆盖度。
- 模块化设计适合不同领域,帮助分析大模型回答的多样性短板。
本文提出ICAT框架,用于衡量长文本生成中多样事实信息的覆盖程度。ICAT将长输出分解为原子性陈述,不仅通过可靠知识源检索验证每条陈述的真实性,还计算这些事实陈述与预期内容维度之间的对齐度。研究了三种不同假设下框架的实现方式,分别对应不同维度可用性和对齐方法的设定。基于TREC Web Track中的多样化任务数据及ClueWeb语料库进行评估,结果表明该框架与人工评判具有强相关性,并对多个主流大模型进行了全面评测。框架支持可解释、细粒度的多样性与覆盖度分析,其模块化设计便于适配不同领域与数据集,是评估大模型生成长文本质量的重要工具。
原文摘要 · Abstract (English)
This paper presents ICAT, an evaluation framework for measuring coverage of diverse factual information in long-form text generation. ICAT breaks down a long output text into a list of atomic claims and not only verifies each claim through retrieval from a (reliable) knowledge source, but also computes the alignment between the atomic factual claims and various aspects expected to be presented in the output. We study three implementations of the ICAT framework, each with a different assumption on the availability of aspects and alignment method. By adopting data from the diversification task in the TREC Web Track and the ClueWeb corpus, we evaluate the ICAT framework. We demonstrate strong correlation with human judgments and provide comprehensive evaluation across multiple state-of-the-art LLMs. Our framework further offers interpretable and fine-grained analysis of diversity and coverage. Its modular design allows for easy adaptation to different domains and datasets, making it a valuable tool for evaluating the qualitative aspects of long-form responses produced by LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。