arXiv:2507.11508cs.CL2025-07EMNLP被引 3

评测酒店亮点摘要的忠实度,发现简单匹配比复杂方法更靠谱。

Real-World Summarization: When Evaluation Reaches Its Limits

  • 用词重合率等简单指标衡量摘要忠实度,效果出人意料好。
  • 人工评估显示,错误信息和不可验证内容是最大风险点。
  • 大模型自评不可靠,易过度或遗漏标注,不适合做评测工具。

我们研究了酒店亮点摘要中对输入数据忠实度的评估问题:即由大模型生成的简短摘要是否准确捕捉住宿特色。通过人类评估实验,包括类别错误判断和跨度级标注,对比传统指标、可训练方法及大模型作为评判者的方法。结果表明,简单的词重合指标在跨领域数据上与人工判断的相关性高达0.63(斯皮尔曼秩相关),常优于复杂方法。尽管大模型能生成高质量摘要,但其自身作为评估者时表现不可靠,易严重低估或高估内容。分析还指出,错误信息和无法验证的信息在实际业务中风险最高。同时,众包评估也面临诸多挑战。

原文摘要 · Abstract (English)

We examine evaluation of faithfulness to input data in the context of hotel highlights: brief LLM-generated summaries that capture unique features of accommodations. Through human evaluation campaigns involving categorical error assessment and span-level annotation, we compare traditional metrics, trainable methods, and LLM-as-a-judge approaches. Our findings reveal that simpler metrics like word overlap correlate surprisingly well with human judgments (Spearman correlation rank of 0.63), often outperforming more complex methods when applied to out-of-domain data. We further demonstrate that while LLMs can generate high-quality highlights, they prove unreliable for evaluation as they tend to severely under- or over-annotate. Our analysis of real-world business impacts shows incorrect and non-checkable information pose the greatest risks. We also highlight challenges in crowdsourced evaluations.

摘要评估忠实度大模型评测酒店数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。