arXiv:2505.04759cs.SEcs.AI2025-05综述被引 2

用ChatGPT零样本分类应用评论,准确率达84.2%。

Exploring Zero-Shot App Review Classification with ChatGPT: Challenges and Potential

  • 直接用ChatGPT零样本分类评论,无需标注数据
  • 在1880条评论上取得0.842的F1分数
  • 适合想快速分析用户反馈的开发者或产品经理

应用评论是用户反馈的重要来源,对应用性能、功能、可用性和整体体验具有关键价值。有效分析这些评论有助于指导应用开发、优先更新功能并提升用户满意度。将评论分类为功能需求与非功能需求,有助于区分具体功能反馈与性能、易用性、可靠性等质量属性的反馈,二者对决策均至关重要。传统方法受限于需大量领域特定标注数据,成本高且耗时。本研究探索使用ChatGPT进行零样本学习,将应用评论分为四类:功能需求、非功能需求、两者皆有或均无。我们在包含1880条手动标注评论的基准数据集上评估其表现,覆盖十款跨领域的应用。结果表明,尽管存在挑战,ChatGPT仍取得了0.842的稳健F1分数。此外,我们分析了评论可读性和长度对分类准确率的影响,并通过人工分析识别出更易误判的类别。

原文摘要 · Abstract (English)

App reviews are a critical source of user feedback, offering valuable insights into an app's performance, features, usability, and overall user experience. Effectively analyzing these reviews is essential for guiding app development, prioritizing feature updates, and enhancing user satisfaction. Classifying reviews into functional and non-functional requirements play a pivotal role in distinguishing feedback related to specific app features (functional requirements) from feedback concerning broader quality attributes, such as performance, usability, and reliability (non-functional requirements). Both categories are integral to informed development decisions. Traditional approaches to classifying app reviews are hindered by the need for large, domain-specific datasets, which are often costly and time-consuming to curate. This study explores the potential of zero-shot learning with ChatGPT for classifying app reviews into four categories: functional requirement, non-functional requirement, both, or neither. We evaluate ChatGPT's performance on a benchmark dataset of 1,880 manually annotated reviews from ten diverse apps spanning multiple domains. Our findings demonstrate that ChatGPT achieves a robust F1 score of 0.842 in review classification, despite certain challenges and limitations. Additionally, we examine how factors such as review readability and length impact classification accuracy and conduct a manual analysis to identify review categories more prone to misclassification.

零样本学习评论分类大模型应用用户反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。