arXiv:2509.14611cs.CL2025-09被引 9

用增强数据提升印尼语电商评论情感分类准确率

Leveraging IndoBERT and DistilBERT for Indonesian Emotion Classification in E-Commerce Reviews

  • 用回译和同义词替换增强数据,提升模型表现
  • IndoBERT经调参后达80%准确率,优于DistilBERT
  • 适合做印尼语NLP、电商情感分析的研究者参考

理解印尼语中的情感对提升电商客户体验至关重要。本研究通过利用先进语言模型IndoBERT和DistilBERT,提升印尼语情绪分类的准确性。关键方法是数据处理,特别是数据增强,包括回译和同义词替换,显著提升了模型性能。经过超参数调优,IndoBERT达到80%的准确率,表明精细数据处理的重要性。虽多模型融合略有提升,但未显著改善效果。结果表明,IndoBERT在印尼语情绪分类中最为有效,数据增强是实现高准确率的关键因素。未来研究应探索其他架构与策略,以提升印尼语NLP任务的泛化能力。

原文摘要 · Abstract (English)

Understanding emotions in the Indonesian language is essential for improving customer experiences in e-commerce. This study focuses on enhancing the accuracy of emotion classification in Indonesian by leveraging advanced language models, IndoBERT and DistilBERT. A key component of our approach was data processing, specifically data augmentation, which included techniques such as back-translation and synonym replacement. These methods played a significant role in boosting the model's performance. After hyperparameter tuning, IndoBERT achieved an accuracy of 80\%, demonstrating the impact of careful data processing. While combining multiple IndoBERT models led to a slight improvement, it did not significantly enhance performance. Our findings indicate that IndoBERT was the most effective model for emotion classification in Indonesian, with data augmentation proving to be a vital factor in achieving high accuracy. Future research should focus on exploring alternative architectures and strategies to improve generalization for Indonesian NLP tasks.

情绪分类IndoBERT数据增强电商文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。