首个可识别豪萨语机器与人类写作的检测工具,准确率达99.2%。
Who Wrote This? Identifying Machine vs Human-Generated Text in Hausa
- 用真实媒体文章和模型生成内容构建豪萨语数据集
- 基于AfroXLMR模型实现99.23%准确率,领先其他模型
- 开源数据集助力低资源语言文本检测研究
大型语言模型(LLMs)在内容生成方面表现优异,但其无监管使用可能引发抄袭和虚假信息传播,尤其对低资源语言构成威胁。现有检测工具多针对英语、法语等高资源语言,本研究首次构建了豪萨语机器生成文本检测系统。我们从七个豪萨语媒体网站收集人类撰写的文本,并利用Gemini-2.0 flash模型根据标题自动生成对应文章。在此基础上,对四种非洲中心预训练模型(AfriTeVa、AfriBERTa、AfroXLMR、AfroXLMR-76L)进行微调,评估结果显示AfroXLMR表现最佳,准确率为99.23%,F1分数为99.21%。研究数据集已公开,以支持后续相关研究。
原文摘要 · Abstract (English)
The advancement of large language models (LLMs) has allowed them to be proficient in various tasks, including content generation. However, their unregulated usage can lead to malicious activities such as plagiarism and generating and spreading fake news, especially for low-resource languages. Most existing machine-generated text detectors are trained on high-resource languages like English, French, etc. In this study, we developed the first large-scale detector that can distinguish between human- and machine-generated content in Hausa. We scrapped seven Hausa-language media outlets for the human-generated text and the Gemini-2.0 flash model to automatically generate the corresponding Hausa-language articles based on the human-generated article headlines. We fine-tuned four pre-trained Afri-centric models (AfriTeVa, AfriBERTa, AfroXLMR, and AfroXLMR-76L) on the resulting dataset and assessed their performance using accuracy and F1-score metrics. AfroXLMR achieved the highest performance with an accuracy of 99.23% and an F1 score of 99.21%, demonstrating its effectiveness for Hausa text detection. Our dataset is made publicly available to enable further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。