arXiv:2605.02374cs.CRcs.CL2026-05

用对抗生成提升少样本机器文本检测的准确率与抗攻击能力

Fight Poison with Poison: Enhancing Robustness in Few-shot Machine-Generated Text Detection with Adversarial Training

论文配图:Fight Poison with Poison: Enhancing Robustness in Few-shot Machine-Generated Text Detection with Adversarial Training
图 1 · 摘自论文原文
  • 用检索增强生成构造类人文本对抗样例,逼迫检测器学习更强特征
  • 在4个数据集上平均F1提升4.95点,对抗攻击成功率降低3.66个百分点
  • 适合需要高鲁棒性的少样本文本检测场景,如舆情监控与内容安全

机器生成文本(MGT)检测对规范网络信息生态至关重要,但现有检测器在少样本设置下性能不佳,且易受对抗性人类化攻击。为在有限标注下构建精准可靠的检测器,本文从攻击者视角出发,在仅输出黑盒环境下分析检测器漏洞。受此启发,提出REACT框架:通过检索增强生成(RAG)驱动的人类化攻击者生成高度类人的对抗样本,目标检测器则利用对比学习从这些对抗样本中稳定学习少样本表征。攻击者与检测器交替更新,实现协同进化。在4个数据集、4种采样量及3组随机种子下的实验表明,相比8个SOTA检测器,REACT平均检测F1提升4.95点,四种强攻击下的平均攻击成功率下降3.66个百分点。

原文摘要 · Abstract (English)

Machine-generated text (MGT) detection is critical for regulating online information ecosystems, yet existing detectors often underperform in few-shot settings and remain vulnerable to adversarial, humanizing attacks. To build accurate and robust detectors under limited supervision, we adopt a threat-modeling perspective and study detector vulnerabilities from an attacker's viewpoint under an output-only black-box setting. Motivated by this perspective, we propose RAG-GuidEd Attacker Strengthens ConTrastive Few-shot Detector (REACT), an adversarial training framework that improves both few-shot detection performance and robustness against attacks. REACT couples a humanization-oriented attacker with a target detector: the attacker leverages retrieval-augmented generation (RAG) to craft highly human-like adversarial examples to evade detection, while the detector learns from these adversaries with a contrastive objective to stabilize few-shot representation learning and enhance robustness. We alternately update the attacker and the detector to enable their co-evolution. Experiments on 4 datasets with 4 shot sizes and 3 random seeds show that REACT improves average detection F1 by 4.95 points over 8 state-of-the-art (SOTA) detectors and reduces the average attack success rate (ASR) under 4 strong attacks by 3.66 percentage points.

文本检测对抗训练少样本学习RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。