arXiv:2510.03502cs.CLcs.AI2025-10

首个大规模阿拉伯语生成文本检测数据集,助力识别大模型伪造内容。

ALHD: A Large-Scale and Multigenre Benchmark Dataset for Arabic LLM-Generated Text Detection

  • 构建涵盖新闻、社交、评论三类文本的40万条阿拉伯语样本
  • 微调BERT模型在检测中表现优于大模型,跨文体泛化能力差
  • 专为防范虚假信息与学术不端设计,适合语言安全研究者

我们提出ALHD,首个专为区分阿拉伯语人类与大模型生成文本而设计的大规模综合数据集。该数据集覆盖新闻、社交媒体、评论三类文体,包含标准语(MSA)与方言,并整合了三个主流大模型生成的超过40万条平衡样本,来源自多个真实人类文本。数据集经过严格预处理、丰富标注及标准化划分,确保可复现性。我们基于此进行基准实验,评估传统分类器、BERT模型及大模型(零样本与少样本)的表现,结果显示微调后的BERT模型性能领先,但跨文体泛化能力不足:尤其在新闻类文本中,大模型生成文本与人类写作风格高度相似,导致检测困难。本研究揭示了当前方法的局限性,也为未来研究指明方向。ALHD为阿拉伯语大模型文本检测及应对虚假信息、学术不端和网络威胁奠定了基础。

原文摘要 · Abstract (English)

We introduce ALHD, the first large-scale comprehensive Arabic dataset explicitly designed to distinguish between human- and LLM-generated texts. ALHD spans three genres (news, social media, reviews), covering both MSA and dialectal Arabic, and contains over 400K balanced samples generated by three leading LLMs and originated from multiple human sources, which enables studying generalizability in Arabic LLM-genearted text detection. We provide rigorous preprocessing, rich annotations, and standardized balanced splits to support reproducibility. In addition, we present, analyze and discuss benchmark experiments using our new dataset, in turn identifying gaps and proposing future research directions. Benchmarking across traditional classifiers, BERT-based models, and LLMs (zero-shot and few-shot) demonstrates that fine-tuned BERT models achieve competitive performance, outperforming LLM-based models. Results are however not always consistent, as we observe challenges when generalizing across genres; indeed, models struggle to generalize when they need to deal with unseen patterns in cross-genre settings, and these challenges are particularly prominent when dealing with news articles, where LLM-generated texts resemble human texts in style, which opens up avenues for future research. ALHD establishes a foundation for research related to Arabic LLM-detection and mitigating risks of misinformation, academic dishonesty, and cyber threats.

文本检测阿拉伯语大模型安全数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。