arXiv:2607.10825cs.CLcs.AI2026-07

用分层采样让大模型用更少词总结出全面公正的用户评价。

Large Language Models for Token-Efficient and Semantic-Preserving Opinion Summarization

  • 先分类再分层采样,选出有代表性的评论子集。
  • 相比传统方法,用1/3以下的词数覆盖更多观点且更平衡。
  • 适合需要高效、公正生成评价摘要的场景。

意见性文本(如商品评论、酒店反馈、社交帖子)蕴含丰富的用户体验与偏好信息,但其规模大、冗余多、分布不均,使得有效分析困难,尤其在生成忠实反映多元观点的摘要时。本文提出一种框架,在基于大语言模型(LLM)的意见摘要中兼顾语义保留与词元效率。通过多维度分类(如情感、主题)结合一系列分层采样策略,预先筛选出紧凑且具代表性的评论子集,再以定制提示引导LLM生成平衡摘要,突出产品或酒店的优缺点。在Amazon商品评论、Tripadvisor酒店评论及X/Twitter帖子上的实验表明,该方法显著降低词元使用量与计算成本,同时在内容覆盖率、平衡性与语义保留方面持续优于传统AI方法与标准LLM摘要基线。

原文摘要 · Abstract (English)

Opinionated text - spanning product reviews, hotel feedback, and social posts - captures rich signals about user experiences, preferences, and concerns. However, the scale, redundancy, and imbalance of such corpora make it challenging to analyze opinions effectively, particularly when the goal is to generate summaries that remain faithful to the diversity of viewpoints expressed. This paper presents a framework that preserves semantics in LLM-based opinion summarization while minimizing token usage. We combine multidimensional classification (e.g., sentiment, topics) with a family of stratified sampling strategies to select compact yet representative subsets of opinions before prompting the LLM. Tailored prompts then produce balanced summaries that surface the salient aspects expressed in the opinions (e.g., strengths and weaknesses of products/hotels). Experiments on Amazon product reviews, Tripadvisor hotel reviews, and X/Twitter posts demonstrate that our method significantly reduces token usage and computational cost while consistently outperforming traditional AI-based and standard LLM summarization baselines in terms of content coverage, balance, and semantic preservation.

意见摘要大模型降本增效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。