拆解评论摘要流程,提升可解释性与实际效果
MOSAIC: Modular Opinion Summarization using Aspect Identification and Clustering
- 分模块设计:主题发现、观点提取、生成摘要
- 比基线覆盖更全观点,且更忠实于原文
- 适合需要可解释性的工业级评论系统
评论是旅行者评估在线商品的核心依据,但现有摘要研究多关注端到端质量,忽视基准可靠性与细粒度洞察的实用性。为此,我们提出MOSAIC——一种面向工业部署的可扩展模块化框架,将摘要任务分解为可解释的组件:主题发现、结构化观点提取和基于证据的摘要生成。通过在真实产品页面上进行在线A/B测试,验证了中间输出的展示能显著改善用户体验,并在完整摘要上线前就带来可衡量价值。离线实验表明,MOSAIC在方面覆盖和忠实度上均优于强基线。关键创新在于引入观点聚类作为系统级组件,在用户评论普遍冗余和噪声高的情况下显著提升忠实度。此外,我们揭示了标准SPACE数据集的可靠性局限,并发布新的开源旅游体验数据集TRECS,以支持更稳健的评估。
原文摘要 · Abstract (English)
Reviews are central to how travelers evaluate products on online marketplaces, yet existing summarization research often emphasizes end-to-end quality while overlooking benchmark reliability and the practical utility of granular insights. To address this, we propose MOSAIC, a scalable, modular framework designed for industrial deployment that decomposes summarization into interpretable components, including theme discovery, structured opinion extraction, and grounded summary generation. We validate the practical impact of our approach through online A/B tests on live product pages, showing that surfacing intermediate outputs improves customer experience and delivers measurable value even prior to full summarization deployment. We further conduct extensive offline experiments to demonstrate that MOSAIC achieves superior aspect coverage and faithfulness compared to strong baselines for summarization. Crucially, we introduce opinion clustering as a system-level component and show that it significantly enhances faithfulness, particularly under the noisy and redundant conditions typical of user reviews. Finally, we identify reliability limitations in the standard SPACE dataset and release a new open-source tour experience dataset (TRECS) to enable more robust evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。