arXiv:2411.07398cs.CLcs.AI2024-11被引 3

用NLI+大模型挖掘应用评论中的隐私担忧,比关键词方法多找出1008条。

Beyond Keywords: A Context-based Hybrid Approach to Mining Ethical Concern-related App Reviews

  • 融合自然语言推理与大模型,理解评论中隐含的伦理关切
  • 在4.3万条心理健康类应用评论中,发现1008条新隐私问题
  • 适合做产品安全评估、伦理审查或用户反馈分析的研究者

随着移动应用在日常生活中普及,伦理问题日益突出。用户常通过应用评论表达对安全、隐私和问责的关注,这些反馈对产品改进至关重要。但涉及伦理的评论使用领域专用术语且词汇多样,自动化提取困难。本文提出一种基于自然语言处理的混合方法,结合自然语言推断(NLI)与解码器型大语言模型(如LLaMA),实现大规模伦理相关评论提取。基于43,647条心理健康类应用评论,研究评估了四种NLI模型以识别潜在隐私评论,并对比了领域特定与通用隐私假设的效果;同时测试四种大模型在分类隐私相关评论上的表现;最终采用最优的NLI与LLM组合进一步挖掘新评论。结果显示,DeBERTa-v3-base-mnli-fever-anli NLI模型搭配领域特定假设效果最佳,而Llama3.1-8B-Instruct在分类任务中表现最优。通过NLI+LLM方法,额外提取出1,008条此前关键词方法未识别的隐私相关评论,验证了该方法的有效性。

原文摘要 · Abstract (English)

With the increasing proliferation of mobile applications in our everyday experiences, the concerns surrounding ethics have surged significantly. Users generally communicate their feedback, report issues, and suggest new functionalities in application (app) reviews, frequently emphasizing safety, privacy, and accountability concerns. Incorporating these reviews is essential to developing successful products. However, app reviews related to ethical concerns generally use domain-specific language and are expressed using a more varied vocabulary. Thus making automated ethical concern-related app review extraction a challenging and time-consuming effort. This study proposes a novel Natural Language Processing (NLP) based approach that combines Natural Language Inference (NLI), which provides a deep comprehension of language nuances, and a decoder-only (LLaMA-like) Large Language Model (LLM) to extract ethical concern-related app reviews at scale. Utilizing 43,647 app reviews from the mental health domain, the proposed methodology 1) Evaluates four NLI models to extract potential privacy reviews and compares the results of domain-specific privacy hypotheses with generic privacy hypotheses; 2) Evaluates four LLMs for classifying app reviews to privacy concerns; and 3) Uses the best NLI and LLM models further to extract new privacy reviews from the dataset. Results show that the DeBERTa-v3-base-mnli-fever-anli NLI model with domain-specific hypotheses yields the best performance, and Llama3.1-8B-Instruct LLM performs best in the classification of app reviews. Then, using NLI+LLM, an additional 1,008 new privacy-related reviews were extracted that were not identified through the keyword-based approach in previous research, thus demonstrating the effectiveness of the proposed approach.

伦理挖掘大模型评论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。