arXiv:2505.23276cs.CLcs.AI2025-05被引 10

首次系统检测阿拉伯语大模型文本,发现可识别的生成指纹。

The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text

  • 通过风格分析对比不同模型与生成方式的阿拉伯语文本特征
  • 在正式场景下检测准确率最高达99.9% F1分数
  • 为阿拉伯语信息真实性的保障提供可落地的检测方案

大型语言模型(LLMs)在生成类人化文本方面取得突破,对教育、社交媒体和学术等领域构成严重信息真实性威胁,尤其在阿拉伯语等低资源语言中更为突出。本文系统研究了阿拉伯语机器生成文本,涵盖标题生成、内容感知生成和文本润色等多种策略,覆盖 ALLaM、Jais、Llama 与 GPT-4 等多种模型架构,在学术与社交媒体两大领域展开分析。风格学分析揭示了人类写作与机器生成阿拉伯语文本间的显著语言模式差异,尽管输出高度拟人,但不同上下文下的模型仍留下可检测的特征痕迹。基于此,我们构建了基于 BERT 的检测模型,在正式语境中达到最高 99.9% 的 F1 分数,并在多模型间保持强鲁棒性。跨领域分析证实了现有泛化挑战。据我们所知,这是迄今最全面的阿拉伯语机器生成文本研究,融合多种提示方法、模型架构与深度风格学分析,为建立语言驱动的稳健检测系统奠定基础,对维护阿拉伯语环境下的信息完整性至关重要。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved unprecedented capabilities in generating human-like text, posing subtle yet significant challenges for information integrity across critical domains, including education, social media, and academia, enabling sophisticated misinformation campaigns, compromising healthcare guidance, and facilitating targeted propaganda. This challenge becomes severe, particularly in under-explored and low-resource languages like Arabic. This paper presents a comprehensive investigation of Arabic machine-generated text, examining multiple generation strategies (generation from the title only, content-aware generation, and text refinement) across diverse model architectures (ALLaM, Jais, Llama, and GPT-4) in academic, and social media domains. Our stylometric analysis reveals distinctive linguistic patterns differentiating human-written from machine-generated Arabic text across these varied contexts. Despite their human-like qualities, we demonstrate that LLMs produce detectable signatures in their Arabic outputs, with domain-specific characteristics that vary significantly between different contexts. Based on these insights, we developed BERT-based detection models that achieved exceptional performance in formal contexts (up to 99.9\% F1-score) with strong precision across model architectures. Our cross-domain analysis confirms generalization challenges previously reported in the literature. To the best of our knowledge, this work represents the most comprehensive investigation of Arabic machine-generated text to date, uniquely combining multiple prompt generation methods, diverse model architectures, and in-depth stylometric analysis across varied textual domains, establishing a foundation for developing robust, linguistically-informed detection systems essential for preserving information integrity in Arabic-language contexts.

语言模型检测技术阿拉伯语信息真实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。