arXiv:2507.04751cs.CL2025-07

用大模型生成融合产品信息的更可信摘要,提升用户决策体验。

LLMs as Architects and Critics for Multi-Source Opinion Summarization

  • 让大模型扮演建筑师和评审员,整合评论与产品参数生成摘要
  • 87%用户偏好新摘要,且与人工判断相关性达0.74
  • 构建首个多维评估数据集,推动该领域研究发展

多源意见摘要(M-OS)在传统意见摘要基础上,引入产品描述、关键特性、规格和评分等元数据,生成融合主观评价与客观属性的综合摘要,有助于用户做出更明智决策。尽管大语言模型(LLMs)在多项自然语言处理任务中表现优异,其在M-OS中的潜力尚未被充分探索。此外,缺乏评估数据集也制约了该领域的进展。为此,我们提出M-OS-EVAL,一个涵盖7个核心维度(流畅性、连贯性、相关性、忠实度、方面覆盖度、情感一致性、具体性)的基准数据集。用户研究表明,平均87%的参与者更偏好M-OS摘要。实验表明,融入事实信息的摘要显著提升用户参与度。尤其值得注意的是,基于提示工程的M-OS-PROMPTS方法在人类评判中表现出更强一致性,平均斯皮尔曼相关系数为ρ = 0.74,优于以往方法。

原文摘要 · Abstract (English)

Multi-source Opinion Summarization (M-OS) extends beyond traditional opinion summarization by incorporating additional sources of product metadata such as descriptions, key features, specifications, and ratings, alongside reviews. This integration results in comprehensive summaries that capture both subjective opinions and objective product attributes essential for informed decision-making. While Large Language Models (LLMs) have shown significant success in various Natural Language Processing (NLP) tasks, their potential in M-OS remains largely unexplored. Additionally, the lack of evaluation datasets for this task has impeded further advancements. To bridge this gap, we introduce M-OS-EVAL, a benchmark dataset for evaluating multi-source opinion summaries across 7 key dimensions: fluency, coherence, relevance, faithfulness, aspect coverage, sentiment consistency, specificity. Our results demonstrate that M-OS significantly enhances user engagement, as evidenced by a user study in which, on average, 87% of participants preferred M-OS over opinion summaries. Our experiments demonstrate that factually enriched summaries enhance user engagement. Notably, M-OS-PROMPTS exhibit stronger alignment with human judgment, achieving an average Spearman correlation of \r{ho} = 0.74, which surpasses the performance of previous methodologies.

意见摘要大模型应用多源融合评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。