警告:现有事实性评估指标不可靠,可能误导模型评价。
Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics
- 在11个数据集上重新测试5种主流事实性指标
- 指标间结果不一致,常错误评估系统性能
- 对改写和远距离引用的输出存在偏差
大型语言模型的进步催生了将其用作自然语言生成评估工具的乐观预期。本文通过在11个摘要、检索增强生成和问答数据集上重新评估五种最先进的事实性指标,挑战这一乐观态度。结果发现,这些评估指标之间存在显著不一致性,且经常错误估计系统级表现,可能导致多种误判。进一步研究表明,这些指标对高度改写的输出以及依赖源文档远端内容的输出存在系统性偏差。我们呼吁使用者在特定领域应用前,务必谨慎对待并手动验证这些指标的可靠性。
原文摘要 · Abstract (English)
Improvements in large language models have led to increasing optimism that they can serve as reliable evaluators of natural language generation outputs. In this paper, we challenge this optimism by thoroughly re-evaluating five state-of-the-art factuality metrics on a collection of 11 datasets for summarization, retrieval-augmented generation, and question answering. We find that these evaluators are inconsistent with each other and often misestimate system-level performance, both of which can lead to a variety of pitfalls. We further show that these metrics exhibit biases against highly paraphrased outputs and outputs that draw upon faraway parts of the source documents. We urge users of these factuality metrics to proceed with caution and manually validate the reliability of these metrics in their domain of interest before proceeding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。