arXiv:2606.00392cs.LGcs.AI2026-06

用约束强化学习让大模型改写文本躲过检测,同时保持原意不变。

Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization

论文配图:Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
图 1 · 摘自论文原文
  • 将文本改写建模为带约束的强化学习问题,确保语义不变
  • 在多个检测器上实现高成功率逃逸,且语义损失控制在指定范围内
  • 对不同检测器、领域和提示都表现稳健,适合对抗性攻击研究

AI文本检测器易受改写攻击,现有方法往往难以精确控制语义保留。直接优化逃逸效果会损害细节语义,而加权奖励设计则依赖敏感权重,难以调控逃逸与语义的平衡。本文将检测器规避型大模型改写问题建模为约束马尔可夫决策过程,以逃逸为目标,语义保留为显式约束。提出检测器规避策略优化(DEPO),一种基于拉格朗日对偶的强化学习算法,采用新型分组策略更新机制,训练中自适应平衡语义保留与逃逸能力。在MAGE、M4、RAID及同行评审数据集上,针对MAGE、RoBERTa、RADAR、Binoculars、Fast-DetectGPT等检测器的实验表明,DEPO在保证语义约束的前提下显著提升逃逸成功率,并展现出跨领域、跨检测器及提示级别的鲁棒性。

原文摘要 · Abstract (English)

AI-text detectors are vulnerable to paraphrasing and detector-guided paraphrasing attacks, but existing detector-evasion methods often lack precise control over semantic preservation. In particular, optimizing directly for detector evasion can degrade fine-grained semantics, whereas scalarized reward designs provide only indirect, weight-sensitive control over the evasion-semantics trade-off. We address this limitation by formulating detector-evasive LLM paraphrasing as a Constrained Markov Decision Process, where detector evasion is the primary objective and semantic preservation is enforced as an explicit constraint. We propose Detector Evasion Policy Optimization (DEPO), a Lagrangian primal-dual reinforcement learning algorithm with a novel GRPO-style group-based policy update. DEPO adaptively balances semantic preservation and detector evasion during training, enabling the policy to improve attack success within a prescribed semantic-preservation region. Experiments on MAGE, M4, RAID, and peer-review datasets, evaluated against MAGE, RoBERTa, RADAR, Binoculars, and Fast-DetectGPT detectors, show that DEPO achieves strong detector evasion while precisely satisfying the semantic preservation constraint. DEPO also exhibits cross-domain, cross-detector, and prompt-level robustness.

大模型改写检测规避强化学习语义保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。