构建首个孟加拉语气候新闻数据集,助力本地化环境话语分析
Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing
- 收集2300篇孟加拉语气候新闻,涵盖政治、科学、立场等多视角
- 提出BanglaBERT-Dhoroni模型,在孟加拉语气候观点识别上表现提升
- 填补低收入国家语言研究空白,适合关注南亚气候传播的研究者
气候变化对全球构成严峻挑战,尤其影响资源匮乏、语言代表性不足的低收入国家。尽管孟加拉国是受气候影响最严重的国家之一,但其孟加拉语相关气候与自然语言处理研究仍存在显著空白。为此,我们提出Dhoroni——一个包含2300篇经标注的孟加拉语气候与环境新闻的数据集,涵盖政治影响、科学数据、真实性、立场检测及利益相关方参与等多个视角。同时,我们开展对Dhoroni的深入探索性分析,并引入基于该数据集微调的BanglaBERT-Dhoroni系列基线模型,用于孟加拉语气候与环境观点检测。本研究显著提升了孟加拉语气候话语的可及性与分析能力,填补了像拥有1.8亿人口的孟加拉国这样气候敏感地区的关键沟通与研究缺口。
原文摘要 · Abstract (English)
Climate change poses critical challenges globally, disproportionately affecting low-income countries that often lack resources and linguistic representation on the international stage. Despite Bangladesh's status as one of the most vulnerable nations to climate impacts, research gaps persist in Bengali-language studies related to climate change and NLP. To address this disparity, we introduce Dhoroni, a novel Bengali (Bangla) climate change and environmental news dataset, comprising a 2300 annotated Bangla news articles, offering multiple perspectives such as political influence, scientific/statistical data, authenticity, stance detection, and stakeholder involvement. Furthermore, we present an in-depth exploratory analysis of Dhoroni and introduce BanglaBERT-Dhoroni family, a novel baseline model family for climate and environmental opinion detection in Bangla, fine-tuned on our dataset. This research contributes significantly to enhancing accessibility and analysis of climate discourse in Bengali (Bangla), addressing crucial communication and research gaps in climate-impacted regions like Bangladesh with 180 million people.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。