提出以读者背景评估摘要信息满足度,弥补现有指标不足
Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation

- 用读者角色和专业背景作为评估信号,捕捉个性化信息需求
- 实验发现主流指标对信息内容变化不敏感,与人类判断不符
- 专家评测表明传统与大模型指标均无法准确衡量信息满足度
摘要评估研究多关注总体质量(如ROUGE、BERTScore)或特定属性(如可读性、事实性),但未能衡量摘要对具体用户的实用性。例如,生物医学研究人员与家庭医生对最新疫苗研究的信息需求不同。查询聚焦摘要虽部分反映需求,但用户常无法在简短查询中完整表达需求。相比之下,读者的背景或身份(角色与专业程度)在不同查询间相对稳定,能有效恢复缺失上下文,是评估摘要是否满足其需求的实用信号。本文评估主流摘要指标对信息差异和身份差异的敏感性,发现包括强大多模态大模型评判指标在内的多数方法,在信息内容扰动测试中表现不佳。进一步开展专家人类评估,基于特定背景和使用场景衡量信息满足度偏好。结果表明,传统及大模型指标均不足以衡量信息满足度,且与人类判断一致性低。
原文摘要 · Abstract (English)
The majority of work on summarization evaluation focuses on general summary quality (e.g., ROUGE, BERTScore) or specific desired properties (e.g., readability, factuality). However, these metrics fail to measure the utility of a summary to an individual user. For example, a biomedical researcher learning about the latest vaccine research will have different informational needs from a family doctor. Query-focused summarization captures part of this need, but in practice, users rarely state everything relevant in a query: a single short query is likely inadequate to distinguish the needs of a researcher from those of a physician. By contrast, a reader's background or persona (their role and expertise) is comparatively stable across queries and recovers much of this missing context, which makes it a practical signal for assessing whether a summary satisfies that reader's needs. In this work, we assess how sensitive popular summarization metrics are to both informational and persona differences, and find that many popular metrics, including strong LLM-as-judge metrics, fail basic perturbation tests of informational content. We additionally conduct an expert human evaluation, measuring summary preferences based on information satisfaction given a specific person's background and use case. We find that both traditional and LLM-based metrics are insufficient measures of information satisfaction and agree poorly with human judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。