arXiv:2506.21686cs.CLcs.LG2025-06被引 4

构建首个孟加拉方言情感分析数据集,助力低资源语言理解

ANUBHUTI: A Comprehensive Corpus For Sentiment Analysis In Bangla Regional Languages

  • 人工将标准孟加拉语译为4种主要方言,确保语义准确
  • 含1万句带多标签情感与主题标注的文本,覆盖政治宗教内容
  • 专家翻译+高一致性校验,适合方言NLP与社会舆情研究

由于语言多样性及标注数据匮乏,孟加拉地区方言的情感分析仍属空白。本文提出ANUBHUTI,一个包含10,000句的手动翻译数据集,将标准孟加拉语转译为莫明辛、诺阿哈利、锡尔赫特和吉大港四种主要方言。数据以政治与宗教内容为主,兼顾中性文本,反映当代孟加拉社会政治面貌。每句采用双重标注:多类主题标签(政治、宗教、中性)与多标签情绪标注(愤怒、轻蔑、厌恶、愉悦、恐惧、悲伤、惊讶)。由母语专家完成翻译与标注,通过Cohen's Kappa检验确保跨方言一致性。数据经系统性检查,剔除缺失、异常与不一致项。ANUBHUTI填补了低资源孟加拉方言情感分析的数据空白,推动更精准、上下文敏感的自然语言处理发展。

原文摘要 · Abstract (English)

Sentiment analysis for regional dialects of Bangla remains an underexplored area due to linguistic diversity and limited annotated data. This paper introduces ANUBHUTI, a comprehensive dataset consisting of 10,000 sentences manually translated from standard Bangla into four major regional dialects Mymensingh, Noakhali, Sylhet, and Chittagong. The dataset predominantly features political and religious content, reflecting the contemporary socio political landscape of Bangladesh, alongside neutral texts to maintain balance. Each sentence is annotated using a dual annotation scheme: multiclass thematic labeling categorizes sentences as Political, Religious, or Neutral, and multilabel emotion annotation assigns one or more emotions from Anger, Contempt, Disgust, Enjoyment, Fear, Sadness, and Surprise. Expert native translators conducted the translation and annotation, with quality assurance performed via Cohens Kappa inter annotator agreement, achieving strong consistency across dialects. The dataset was further refined through systematic checks for missing data, anomalies, and inconsistencies. ANUBHUTI fills a critical gap in resources for sentiment analysis in low resource Bangla dialects, enabling more accurate and context aware natural language processing.

情感分析方言识别低资源语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。