arXiv:2509.00285cs.CLcs.AI2025-09被引 4

针对海量评论生成个性化观点摘要,提升总结精准度与可读性。

OpinioRAG: Towards Generating User-Centric Opinion Highlights from Large-scale Online Reviews

  • 基于检索增强生成框架,无需训练即可实现高效个性化摘要。
  • 在超千条评论实体上验证,生成摘要与专家标注高度一致。
  • 提出无参考的语义一致性评估指标,适合情感密集型内容评价。

我们研究从大规模用户评论中生成观点摘要的问题,现有方法或难以扩展,或生成通用化、缺乏个性化的总结。为此,我们提出 OpinioRAG——一种可扩展、无需训练的框架,结合基于 RAG 的证据检索与大语言模型,高效生成定制化摘要。此外,我们设计了适用于情感丰富领域的新型无参考验证指标,能够细粒度、上下文敏感地评估事实一致性。为支持评估,我们构建了首个包含上千条评论/实体的长篇评论数据集,配有无偏专家摘要和人工标注查询。大量实验揭示关键挑战,提供可操作改进方案,推动未来研究发展,并确立 OpinioRAG 在规模化生成准确、相关且结构化摘要方面的鲁棒性。

原文摘要 · Abstract (English)

We study the problem of opinion highlights generation from large volumes of user reviews, often exceeding thousands per entity, where existing methods either fail to scale or produce generic, one-size-fits-all summaries that overlook personalized needs. To tackle this, we introduce OpinioRAG, a scalable, training-free framework that combines RAG-based evidence retrieval with LLMs to efficiently produce tailored summaries. Additionally, we propose novel reference-free verification metrics designed for sentiment-rich domains, where accurately capturing opinions and sentiment alignment is essential. These metrics offer a fine-grained, context-sensitive assessment of factual consistency. To facilitate evaluation, we contribute the first large-scale dataset of long-form user reviews, comprising entities with over a thousand reviews each, paired with unbiased expert summaries and manually annotated queries. Through extensive experiments, we identify key challenges, provide actionable insights into improving systems, pave the way for future research, and position OpinioRAG as a robust framework for generating accurate, relevant, and structured summaries at scale.

观点摘要RAG评论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。