arXiv:2411.04093cs.CL2024-11被引 4

构建多视角政治文本摘要数据集,评估模型忠实还原立场的能力。

Summarization of Opinionated Political Documents with Varied Perspectives

  • 设计新任务与数据集,独立摘要不同政治立场的新闻段落。
  • 11个模型在忠实还原立场上表现普遍不佳,尤其大模型也存在偏差。
  • 分析输入文本特征如何影响摘要提取行为,揭示模型局限性。

全球党派对立和极化现象加剧,尤其在总统选举期间更为明显。能够生成多元观点准确摘要的模型,有助于通过暴露用户于不同立场来缓解极化。本文提出一种新数据集和任务,旨在对来自观点性新闻文章的文本段落独立生成各政治立场的摘要。为此,我们构建了一个多维度评估框架,对11种不同规模和架构的摘要模型及大语言模型进行了自动与人工评估。尽管近期模型如GPT-4o表现较好,但所有模型在忠实再现目标立场方面仍面临挑战。我们的分析聚焦于输入文档特征如何影响摘要的提取行为,揭示了模型在立场保持上的系统性缺陷。

原文摘要 · Abstract (English)

Global partisan hostility and polarization has increased, and this polarization is heightened around presidential elections. Models capable of generating accurate summaries of diverse perspectives can help reduce such polarization by exposing users to alternative perspectives. In this work, we introduce a novel dataset and task for independently summarizing each political perspective in a set of passages from opinionated news articles. For this task, we propose a framework for evaluating different dimensions of perspective summary performance. We benchmark 11 summarization models and LLMs of varying sizes and architectures through both automatic and human evaluation. While recent models like GPT-4o perform well on this task, we find that all models struggle to generate summaries that are faithful to the intended perspective. Our analysis of summaries focuses on how extraction behavior is impacted by features of the input documents.

文本摘要政治极化多视角大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。