arXiv:2509.20953cs.AI2025-09中稿 · presentation at AC…综述

用大模型让评分与评论更匹配,挖掘用户真实反馈。

Beyond Stars: Bridging the Gap Between Ratings and Review Sentiment with LLM

  • 用大模型+结构化提示,理解评论中的语境和讽刺
  • 在三个数据集上准确率显著高于传统方法
  • 适合做产品优化的团队或需要深度洞察的分析师

我们提出一种先进的移动端应用评论分析方法,以克服传统星级评分系统的局限性。星级评分虽直观易用,却难以捕捉详细评论中的细微反馈。传统NLP技术(如词典法和经典机器学习分类器)在理解上下文、领域术语及反讽等语言特征方面表现不佳。为此,我们设计了一种模块化框架,利用大语言模型(LLM)并结合结构化提示技术,量化评分与文本情感之间的差异,提取特征级洞察,并通过检索增强型对话问答(RAG-QA)支持评论的交互式探索。在AWARE、Google Play和Spotify三个多样化数据集上的全面实验表明,该方法在复杂多变的评论场景中显著优于基线模型,提升了准确性、鲁棒性与可操作洞察力。

原文摘要 · Abstract (English)

We present an advanced approach to mobile app review analysis aimed at addressing limitations inherent in traditional star-rating systems. Star ratings, although intuitive and popular among users, often fail to capture the nuanced feedback present in detailed review texts. Traditional NLP techniques -- such as lexicon-based methods and classical machine learning classifiers -- struggle to interpret contextual nuances, domain-specific terminology, and subtle linguistic features like sarcasm. To overcome these limitations, we propose a modular framework leveraging large language models (LLMs) enhanced by structured prompting techniques. Our method quantifies discrepancies between numerical ratings and textual sentiment, extracts detailed, feature-level insights, and supports interactive exploration of reviews through retrieval-augmented conversational question answering (RAG-QA). Comprehensive experiments conducted on three diverse datasets (AWARE, Google Play, and Spotify) demonstrate that our LLM-driven approach significantly surpasses baseline methods, yielding improved accuracy, robustness, and actionable insights in challenging and context-rich review scenarios.

大模型应用情感分析用户反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。