arXiv:2505.24523cs.CLcs.AI2025-05ACL被引 11

用风格迁移让生成文本骗过检测器,暴露现有方法漏洞

Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors

  • 用直接偏好优化微调模型,让生成文本更像真人写作
  • 少量样本就能让检测器准确率大幅下降,最高超80%失效
  • 适合关注检测鲁棒性与对抗攻击的研究者

近期生成式AI和大语言模型的发展使得合成内容高度逼真,引发虚假信息与操纵的担忧。然而,机器生成文本检测仍面临挑战,主要因为缺乏能评估真实场景泛化能力的稳健基准。本文提出一个测试框架,用于评估主流检测器(如Mage、Radar、LLM-DetectAIve)对语义层面对抗攻击的鲁棒性。通过使用直接偏好优化(DPO)微调语言模型,使生成文本的写作风格向真人文本靠拢,从而利用检测器对风格特征的依赖,使其难以识别。我们还分析了对齐过程带来的语言变化及检测器所依赖的关键特征。结果表明,仅需少量样本即可显著降低检测性能,部分检测器失效率达80%以上,凸显提升检测方法鲁棒性、应对未见领域文本的重要性。

原文摘要 · Abstract (English)

Recent advancements in Generative AI and Large Language Models (LLMs) have enabled the creation of highly realistic synthetic content, raising concerns about the potential for malicious use, such as misinformation and manipulation. Moreover, detecting Machine-Generated Text (MGT) remains challenging due to the lack of robust benchmarks that assess generalization to real-world scenarios. In this work, we present a pipeline to test the resilience of state-of-the-art MGT detectors (e.g., Mage, Radar, LLM-DetectAIve) to linguistically informed adversarial attacks. To challenge the detectors, we fine-tune language models using Direct Preference Optimization (DPO) to shift the MGT style toward human-written text (HWT). This exploits the detectors' reliance on stylistic clues, making new generations more challenging to detect. Additionally, we analyze the linguistic shifts induced by the alignment and which features are used by detectors to detect MGT texts. Our results show that detectors can be easily fooled with relatively few examples, resulting in a significant drop in detection performance. This highlights the importance of improving detection methods and making them robust to unseen in-domain texts.

文本检测对抗攻击风格迁移大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。