arXiv:2410.13037cs.CLcs.AI2024-10被引 1

用大模型总结超长用户评论,更准更忠实。

LFOSum: Summarizing Long-form Opinions with Large Language Models

  • 不需训练,直接用大模型处理千条以上长评论
  • 新数据集含专家标注的精准摘要,支持严格评估
  • 新增无参考评价指标,更细致判断摘要是否靠谱

在线评论在购物、酒店、餐饮等多个领域对消费者决策起关键作用。然而,评论数量庞大且常重复或无关,导致信息过载,用户难以提取有效信息。传统意见摘要模型难以处理长文本和大量评论,而新出现的大语言模型方法往往生成不准确或不忠实的摘要。为此,本文提出:(1) 一个包含上千条评论的长篇用户评论新数据集;(2) 两种无需训练的基于大模型的摘要方法,可扩展至长输入;(3) 自动化评估指标。该数据集配有领域专家撰写的深度、中立批判性摘要,作为评估基准。此外,新提出的无参考评估指标能更精细、情境敏感地衡量摘要的忠实度。我们在多种开源与闭源大模型上进行了基准测试。评估显示,大模型在长篇摘要中仍难以平衡情感表达与格式一致性,但当相关信息被聚焦检索时,开源模型可缩小差距。

原文摘要 · Abstract (English)

Online reviews play a pivotal role in influencing consumer decisions across various domains, from purchasing products to selecting hotels or restaurants. However, the sheer volume of reviews -- often containing repetitive or irrelevant content -- leads to information overload, making it challenging for users to extract meaningful insights. Traditional opinion summarization models face challenges in handling long inputs and large volumes of reviews, while newer Large Language Model (LLM) approaches often fail to generate accurate and faithful summaries. To address those challenges, this paper introduces (1) a new dataset of long-form user reviews, each entity comprising over a thousand reviews, (2) two training-free LLM-based summarization approaches that scale to long inputs, and (3) automatic evaluation metrics. Our dataset of user reviews is paired with in-depth and unbiased critical summaries by domain experts, serving as a reference for evaluation. Additionally, our novel reference-free evaluation metrics provide a more granular, context-sensitive assessment of summary faithfulness. We benchmark several open-source and closed-source LLMs using our methods. Our evaluation reveals that LLMs still face challenges in balancing sentiment and format adherence in long-form summaries, though open-source models can narrow the gap when relevant information is retrieved in a focused manner.

大模型评论摘要长文本评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。