用大数据框架提升亚马逊评论欺诈检测准确率
Leveraging Big Data Frameworks for Spam Detection in Amazon Reviews
- 基于大规模数据框架分析评论特征,识别虚假行为
- 逻辑回归模型达到90.35%准确率,有效提升检测能力
- 适合电商风控、平台审核人员参考使用
在数字时代,在线购物已成为日常生活的一部分。产品评论显著影响消费者购买决策并建立买家信任。然而,虚假评论的泛滥破坏了这种信任,可能误导消费者并损害卖家声誉。本研究通过在大规模亚马逊产品评论数据集上应用先进的大数据分析与机器学习方法,解决这一紧迫问题。目标是准确检测和分类垃圾评论,以增强评论的真实性。利用可扩展的大数据框架,我们高效处理和分析海量评论数据,提取出反映欺诈行为的关键特征。研究表明,多种机器学习分类器在检测垃圾评论方面具有有效性,其中逻辑回归模型准确率达到90.35%,有助于构建更可信、透明的在线购物环境。
原文摘要 · Abstract (English)
In this digital era, online shopping is common practice in our daily lives. Product reviews significantly influence consumer buying behavior and help establish buyer trust. However, the prevalence of fraudulent reviews undermines this trust by potentially misleading consumers and damaging the reputations of the sellers. This research addresses this pressing issue by employing advanced big data analytics and machine learning approaches on a substantial dataset of Amazon product reviews. The primary objective is to detect and classify spam reviews accurately so that it enhances the authenticity of the review. Using a scalable big data framework, we efficiently process and analyze a large scale of review data, extracting key features indicative of fraudulent behavior. Our study illustrates the utility of various machine learning classifiers in detecting spam reviews, with Logistic Regression achieving an accuracy of 90.35%, thus contributing to a more trustworthy and transparent online shopping environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。