构建4.5万条阿尔及利亚方言评论数据集,助力阿拉伯语方言情感分析研究。
Algerian Dialect
- 收集30+媒体频道的YouTube评论,人工标注五类情感标签。
- 含时间戳、点赞数等元数据,覆盖4.5万条阿尔及利亚阿拉伯语评论。
- 填补阿拉伯方言资源空白,适合做社会媒体分析与方言NLP研究。
我们提出了阿尔及利亚方言数据集,包含45,000条来自阿尔及利亚新闻与媒体频道的YouTube评论,均用阿尔及利亚阿拉伯语书写。评论通过YouTube Data API从30多个媒体来源收集,每条评论由人工标注为五类情感:非常负面、负面、中性、正面、非常正面。数据集还包含丰富的元数据,如采集时间戳、点赞数、视频链接和标注日期。该数据集旨在解决阿尔及利亚方言公开资源稀缺的问题,支持情感分析、方言阿拉伯语NLP及社交媒体分析研究。数据集已通过Mendeley Data公开发布,采用CC BY 4.0许可,访问地址:https://doi.org/10.17632/zzwg3nnhsz.2。
原文摘要 · Abstract (English)
We present Algerian Dialect, a large-scale sentiment-annotated dataset consisting of 45,000 YouTube comments written in Algerian Arabic dialect. The comments were collected from more than 30 Algerian press and media channels using the YouTube Data API. Each comment is manually annotated into one of five sentiment categories: very negative, negative, neutral, positive, and very positive. In addition to sentiment labels, the dataset includes rich metadata such as collection timestamps, like counts, video URLs, and annotation dates. This dataset addresses the scarcity of publicly available resources for Algerian dialect and aims to support research in sentiment analysis, dialectal Arabic NLP, and social media analytics. The dataset is publicly available on Mendeley Data under a CC BY 4.0 license at https://doi.org/10.17632/zzwg3nnhsz.2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。