构建首个阿尔及利亚方言虚假新闻与情感分析语料库,填补低资源语言空白。
FASSILA: A Corpus for Algerian Dialect Fake News Detection and Sentiment Analysis
- 构建包含10,087句、19,497个唯一词的阿尔及利亚方言语料库
- 跨7个领域标注,人工标注一致性高,质量可靠
- 开源共享,助力低资源语言自然语言处理研究
在低资源语言背景下,阿尔及利亚方言(AD)因缺乏标注语料库而面临处理难题,尤其影响依赖语料训练与评估的机器学习应用。本研究提出名为FASSILA的专用语料库,用于阿尔及利亚方言的虚假新闻检测与情感分析。该语料库共包含10,087条句子,涵盖超过19,497个独特词汇,覆盖七个不同领域。研究设计了虚假新闻检测与情感分析的标注方案,并详细描述数据收集、清洗与标注流程。显著的标注者间一致性表明该标注方案能产生高质量、一致的标注结果。随后使用基于BERT的模型与传统机器学习模型进行分类实验,取得良好效果,揭示了未来研究方向。该数据集已开源至GitHub(https://github.com/amincoding/FASSILA),以推动该领域发展。
原文摘要 · Abstract (English)
In the context of low-resource languages, the Algerian dialect (AD) faces challenges due to the absence of annotated corpora, hindering its effective processing, notably in Machine Learning (ML) applications reliant on corpora for training and assessment. This study outlines the development process of a specialized corpus for Fake News (FN) detection and sentiment analysis (SA) in AD called FASSILA. This corpus comprises 10,087 sentences, encompassing over 19,497 unique words in AD, and addresses the significant lack of linguistic resources in the language and covers seven distinct domains. We propose an annotation scheme for FN detection and SA, detailing the data collection, cleaning, and labelling process. Remarkable Inter-Annotator Agreement indicates that the annotation scheme produces consistent annotations of high quality. Subsequent classification experiments using BERT-based models and ML models are presented, demonstrate promising results and highlight avenues for further research. The dataset is made freely available on GitHub (https://github.com/amincoding/FASSILA) to facilitate future advancements in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。