构建多语言长文本问答评估基准,揭示现有自动评估方法不足
LFQA-E: Carefully Benchmarking Long-form QA Evaluation
- 构建包含1618个问题、7323组对比的多语言评估基准
- 17种评估方法均无法接近人工判断水平,难以捕捉长回答信息密度
- 适合研究自动评估与长文本问答的学者使用
长文本问答(LFQA)要求对开放性问题生成详尽的段落级回答,其评估因信息丰富性和响应格式灵活而极具挑战。现有评估基准常缺乏参考答案,规模与主题覆盖有限,可靠性不足。为此,我们提出LFQA-E,一个结构严谨、多语言、基于参考答案的基准,用于严格评估LFQA的自动评估指标。该基准涵盖15个主题,共1618个问题和7323组成对比较,数据来源包括在线查询与考试题。我们测试了五类共17种评估指标,结果表明:所有现有自动指标均无法媲美人工判断,暴露出其在捕捉长回答密集信息方面的局限。此外,我们深入分析失败案例与指标泛化能力,为未来评估方法发展提供指导。基准与代码已公开于https://github.com/YuchenFan48/LFQA-E。
原文摘要 · Abstract (English)
Long-Form Question Answering (LFQA) involves generating comprehensive, paragraph-level responses to open-ended questions, which poses a significant challenge for evaluation due to the richness of information and flexible response format. Existing LFQA-evaluation benchmarks often lack reference answers and are limited in size and topic coverage, reducing their reliability. To address this gap, we introduce LFQA-E, a well-constructed, multilingual, and reference-based benchmark designed to rigorously evaluate automatic metrics for LFQA. LFQA-E comprises 1618 questions and 7323 pairwise comparisons across 15 topics, drawn from diverse sources such as online queries and examination questions, thereby enabling a comprehensive assessment of evaluation metrics. We examine five categories of metrics, encompassing 17 specific methods, using LFQA-E. The results demonstrate that none of the existing automatic metrics perform comparably to human judgments, highlighting their inability to capture the dense information in long-form responses. Furthermore, we present a detailed analysis of the failure cases and the generalization capacity of these metrics, offering insights to guide the future development of LFQA evaluation methods. The benchmark and code are available at https://github.com/YuchenFan48/LFQA-E.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。