测试发现多数事实性评估指标易被操纵,可靠性存疑。
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
- 用浅层分类器区分易难案例,检验指标实际能力
- 所有指标在需要深度推理的难题上表现显著下降
- 添加无关句子即可虚增分数,仅提示式ChatGPT-DA较稳健
现代大语言模型可生成高度可读的摘要,传统评估指标如ROUGE已饱和。但模型常引入与源文本不一致或无支持的信息,自动检测此类微妙事实错误仍具挑战。为此发展出多种事实一致性度量方法,但它们是否真正衡量了事实性?本研究对多种自动事实性度量(包括专用模型和基于LLM的提示方法)进行压力测试。通过浅层分类器将事实评估样本分为仅依赖表面特征的“简单”例与需深层推理的“困难”例,发现所有指标在困难例上性能大幅下降。此外,部分指标对无害的改写更敏感,甚至比对事实修正更敏感。进一步实验表明,多数指标可通过添加无意义内容句子被人为提高分数。其中,基于提示的ChatGPT-DA方法最稳健,但其判断可能过度依赖模型参数知识而非给定参考文本。整体结果质疑当前事实性度量的可靠性,并引发对这些指标真实测量目标的反思。
原文摘要 · Abstract (English)
Modern LLMs can now produce highly readable abstractive summaries, to the point that traditional automated metrics for evaluating summary quality, such as ROUGE, have saturated. However, LLMs still sometimes introduce inaccuracies into summaries, i.e., information inconsistent with or unsupported by the corresponding source. Measuring the occurrence of these often subtle factual inconsistencies automatically has proved challenging. This in turn has motivated development of metrics intended to measure the factual consistency of generated summaries against sources. But are these approaches measuring what they purport to? Or are they mostly exploiting artifacts? In this work, we stress test a range of automatic factuality metrics, including specialized models and LLM-based prompting methods, to probe what they actually capture. Using a shallow classifier to separate ``easy'' examples for factual evaluation where surface features suffice from ``hard'' cases requiring deeper reasoning, we find that all metrics show substantial performance drops on the latter. Furthermore, some metrics are more sensitive to benign, fact-preserving edits than to factual corrections. Building on this observation, we demonstrate that most automatic factuality metrics can be gamed, i.e., their scores can be artificially inflated by appending innocuous, content-free sentences to summaries. Among the metrics tested, the prompt based ChatGPT-DA approach is the most robust and reliable. However, this comes with a notable caveat: Prompting LLMs to assess factuality may overly rely on their parametric knowledge rather than the provided reference when making judgments. Taken together, our findings call into question the reliability of current factuality metrics and prompt a broader reflection on what these metrics are truly measuring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。