用XGBoost分析印尼电商视频评论,预测用户满意度。
Sentiment Analysis and Customer Satisfaction Prediction on E-Commerce Platforms Based on YouTube Comments Using the XGBoost Algorithm

- 结合TF-IDF与XGBoost构建评论情感与满意度预测模型。
- 经PyCaret优化后,模型在分类任务中表现更稳定。
- 发现电商评论中大量使用政治术语影响用户情绪判断。
印度尼西亚数字商业的指数级扩张显著推动消费者互动向以视频为核心的社交网络(如YouTube)转移。由此带来的海量非结构化、多语境评论,给人工情感追踪带来巨大挑战。本研究基于从电商评测视频中获取的二次数据集,利用极端梯度提升(XGBoost)架构与词频-逆文档频率(TF-IDF)向量化方法,构建客户满意度预测模型。原始文本经过严格预处理,生成标准化数值特征。实验结果表明,经PyCaret优化的机器学习框架在分类任务中展现出更强的鲁棒性。除常规性能指标外,词汇分析与特征重要性映射揭示:电商话语中广泛渗透社会政治术语,最终影响观众满意度极性。
原文摘要 · Abstract (English)
The exponential expansion of digital commerce in Indonesia has significantly shifted consumer interactions toward video-centric social networks, particularly YouTube. Consequently, the sheer volume of unstructured, multi-contextual comments poses a tremendous challenge for manual sentiment tracking. This study investigates and constructs a predictive model for customer satisfaction leveraging the Extreme Gradient Boosting (XGBoost) architecture coupled with Term Frequency-Inverse Document Frequency (TF-IDF) vectorization. By utilizing a secondary dataset of YouTube comments retrieved from e-commerce review videos, the raw text underwent rigorous preprocessing to generate normalized numerical features. The experimental results demonstrate that the PyCaret-optimized machine learning framework delivers superior classification resilience. Beyond standard performance metrics, lexical evaluations and feature-importance mapping uncover a notable phenomenon: e-commerce discourse is heavily infiltrated by socio-political terminologies, which ultimately influence the polarity of audience satisfaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。