对比扩散模型与自回归模型生成文本,发现现有检测方法失效
Can You Detect the Difference?
- 用2000样本系统比较扩散与自回归生成文本特征
- 扩散生成文本在困惑度和突发性上接近人类写作
- 需开发针对扩散模型的新型检测方法
大语言模型的快速发展引发了对人工智能生成文本可检测性的担忧。虽然风格度量在自回归输出中表现良好,但其在扩散模型上的效果尚不明确。我们首次系统性地比较了基于扩散的文本(LLaDA)与自回归文本(LLaMA),使用2000个样本。困惑度、突发性、词汇多样性、可读性以及BLEU/ROUGE得分显示,LLaDA在困惑度和突发性上接近人类文本,导致面向自回归模型的检测器产生高假阴性率。而LLaMA虽有更低困惑度,但词汇保真度下降。单一指标无法区分扩散生成文本与人类写作。研究强调需要开发扩散感知检测器,并提出混合模型、扩散特异性风格特征和鲁棒水印等方向。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has raised concerns about reliably detecting AI-generated text. Stylometric metrics work well on autoregressive (AR) outputs, but their effectiveness on diffusion-based models is unknown. We present the first systematic comparison of diffusion-generated text (LLaDA) and AR-generated text (LLaMA) using 2 000 samples. Perplexity, burstiness, lexical diversity, readability, and BLEU/ROUGE scores show that LLaDA closely mimics human text in perplexity and burstiness, yielding high false-negative rates for AR-oriented detectors. LLaMA shows much lower perplexity but reduced lexical fidelity. Relying on any single metric fails to separate diffusion outputs from human writing. We highlight the need for diffusion-aware detectors and outline directions such as hybrid models, diffusion-specific stylometric signatures, and robust watermarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。