arXiv:2504.01132cs.CLcs.AI2025-04EMNLP被引 2

用改写程度量化叙事判断的主观性,提升评估可靠性。

Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding

  • 用大模型生成摘要改写,衡量观点模糊程度
  • 改写量越大,说明原判断越主观,相关性提升21%
  • 适合需要细致评估叙事真实性的研究者

判断一个陈述是否忠实于源文档是多个领域的重要问题。传统方法通常将该任务视为二元判断:支持或不支持。但在许多情况下,陈述的支持与否存在歧义,例如需基于证据进行推断,不同人可能合理得出相反结论。强制二元标签会降低评估可靠性。本文重新定义该任务,以应对模糊陈述中的主观性。提出使用大模型生成的摘要改写作为细粒度评估手段:陈述需修改多少才能消除歧义?改写程度和修改量可自动转化为‘模糊改写度量’(ARM),提供比二元判断更丰富的反馈信号。聚焦叙事摘要领域,因其尤其容易产生歧义和主观解读。实验表明,采用ARM后,标注者对陈述真实性的意见一致性提升了21个百分点,证明主观性显著降低。

原文摘要 · Abstract (English)

Determining faithfulness of a claim to a source document is an important problem across many domains. This task is generally treated as a binary judgment of whether the claim is supported or unsupported in relation to the source. In many cases, though, whether a claim is supported can be ambiguous. For instance, it may depend on making inferences from given evidence, and different people can reasonably interpret the claim as either supported or unsupported based on their agreement with those inferences. Forcing binary labels upon such claims lowers the reliability of evaluation. In this work, we reframe the task to manage the subjectivity involved with factuality judgments of ambiguous claims. We introduce LLM-generated edits of summaries as a method of providing a nuanced evaluation of claims: how much does a summary need to be edited to be unambiguous? Whether a claim gets rewritten and how much it changes can be used as an automatic evaluation metric, the Ambiguity Rewrite Metric (ARM), with a much richer feedback signal than a binary judgment of faithfulness. We focus on the area of narrative summarization as it is particularly rife with ambiguity and subjective interpretation. We show that ARM produces a 21% absolute improvement in annotator agreement on claim faithfulness, indicating that subjectivity is reduced.

主观性评估叙事理解大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。