arXiv:2412.13161cs.CLcs.CV2024-12被引 7

构建首个大规模孟加拉语-英语混合电商评论数据集,助力多语言情感分析。

BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce

  • 收集174万条跨语言电商评论,涵盖320万条评分与12.8万种商品。
  • 基于该数据集训练的模型在情感分析任务中达到94%准确率和0.94 F1值。
  • 适合研究低资源语言、混合语态自然语言处理及电商评论分析的学者。

本文提出BanglishRev数据集,是目前针对孟加拉语、英语及两者混合形式(即孟加拉语用拉丁字母书写,称作Banglish)的电商评论最大规模数据集。数据来自面向孟加拉语人群的在线电商平台,包含174万条评论、320万条评分信息,覆盖12.8万种商品。每条评论附有丰富元数据,包括评分、发布时间、购买时间、点赞/点踩数、卖家回复、图片等。为验证其在情感分析中的有效性,以评分大于3为正向、小于等于3为负向,训练了BanglishBERT模型。在已有的手动标注混合语种评论数据集上测试,模型取得94%准确率和0.94 F1分数,显著优于现有方法。文中还探讨了数据集中观察到的语言模式与未来研究方向。数据集可通过https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset获取。

原文摘要 · Abstract (English)

This work presents the BanglishRev Dataset, the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online e-commerce platforms targeting the Bengali population. It includes an extensive array of related metadata for each of the reviews including the rating given by the reviewer, date the review was posted and date of purchase, number of likes, dislikes, response from the seller, images associated with the review etc. With sentiment analysis being the most prominent usage of review datasets, experimentation with a binary sentiment analysis model with the review rating serving as an indicator of positive or negative sentiment was conducted to evaluate the effectiveness of the large amount of data presented in BanglishRev for sentiment analysis tasks. A BanglishBERT model is trained on the data from BanglishRev with reviews being considered labeled positive if the rating is greater than 3 and negative if the rating is less than or equal to 3. The model is evaluated by being testing against a previously published manually annotated dataset for e-commerce reviews written in a mixture of Bangla, English and Banglish. The experimental model achieved an exceptional accuracy of 94\% and F1 score of 0.94, demonstrating the dataset's efficacy for sentiment analysis. Some of the intriguing patterns and observations seen within the dataset and future research directions where the dataset can be utilized is also discussed and explored. The dataset can be accessed through https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.

多语言情感分析电商评论低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。