arXiv:2506.10785cs.SEcs.AI2025-06被引 3

分析89万条用户评论,揭示人们对AI手机应用的真实评价与痛点。

What Users Value and Critique: Large-Scale Analysis of User Feedback on AI-Powered Mobile Apps

  • 构建多阶段分析流程,精准识别用户反馈中的情感与主题。
  • 提取超百万个细粒度情感-主题对,发现18个正面与15个负面话题。
  • 揭示用户既夸功能又吐槽缺陷的复杂体验,适合产品优化参考。

人工智能(AI)功能在各类手机应用中迅速普及,涵盖生产力、教育、娱乐和创意等领域。然而,用户如何感知、评估和批评这些功能仍不明确,主要受限于用户反馈量过大。本文首次开展大规模研究,基于涵盖14个类别、292个AI驱动应用的精选数据集,分析了来自Google Play的89.4万条专门针对AI功能的评论。我们开发并验证了一个多阶段分析流程,从人工标注基准开始,系统评估大语言模型(LLMs)及提示策略。各阶段包括评论分类、方面-情感提取和聚类,均通过准确性与一致性验证。该流程实现高精度、可扩展的用户反馈分析,共提取超过一百万个方面-情感对,并聚类为18个正面与15个负面用户话题。分析显示,用户关注主题有限:正面评价聚焦于效率、可靠性和个性化服务;负面反馈则集中在技术故障(如扫描与识别)、价格问题及语言支持不足。同一评论中常同时出现对某功能满意与对另一功能不满的复杂情感,传统方法因孤立处理正负反馈或依赖粗粒度分析而易忽略此类细节。本方法更真实反映真实用户体验。类别感知分析进一步揭示通用满意度驱动因素与领域特异性困扰。

原文摘要 · Abstract (English)

Artificial Intelligence (AI)-powered features have rapidly proliferated across mobile apps in various domains, including productivity, education, entertainment, and creativity. However, how users perceive, evaluate, and critique these AI features remains largely unexplored, primarily due to the overwhelming volume of user feedback. In this work, we present the first comprehensive, large-scale study of user feedback on AI-powered mobile apps, leveraging a curated dataset of 292 AI-driven apps across 14 categories with 894K AI-specific reviews from Google Play. We develop and validate a multi-stage analysis pipeline that begins with a human-labeled benchmark and systematically evaluates large language models (LLMs) and prompting strategies. Each stage, including review classification, aspect-sentiment extraction, and clustering, is validated for accuracy and consistency. Our pipeline enables scalable, high-precision analysis of user feedback, extracting over one million aspect-sentiment pairs clustered into 18 positive and 15 negative user topics. Our analysis reveals that users consistently focus on a narrow set of themes: positive comments emphasize productivity, reliability, and personalized assistance, while negative feedback highlights technical failures (e.g., scanning and recognition), pricing concerns, and limitations in language support. Our pipeline surfaces both satisfaction with one feature and frustration with another within the same review. These fine-grained, co-occurring sentiments are often missed by traditional approaches that treat positive and negative feedback in isolation or rely on coarse-grained analysis. To this end, our approach provides a more faithful reflection of the real-world user experiences with AI-powered apps. Category-aware analysis further uncovers both universal drivers of satisfaction and domain-specific frustrations.

用户研究AI应用情感分析反馈挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。