攻击机器文本检测器仍留有风格痕迹,多文档分析才是可靠检测关键。
Attacks on Machine-Text Detectors Retain Stylistic Fingerprints
- 提出新改写方法,同时优化隐蔽性与人类写作风格。
- 单文档检测易被绕过,但多文档分析能重新识别机器文本。
- 适合关注文本检测安全性和对抗攻击的研究者。
尽管机器文本检测技术取得进展,但文本易被操纵以逃避检测,暗示问题可能本质难解。本文研究此类规避策略的极限,发现当前攻击(从提示工程到检测器引导优化)虽能显著降低标准检测器性能,却无法消除机器文本的潜在风格特征。我们证明,利用风格特征空间的少样本检测器对这些规避手段具有鲁棒性,能可靠识别经专门调优以避免检测的样本。然而,我们引入一种新型改写方法,在优化不可检测性的同时保持特定人类风格,结果表明该攻击可有效绕过所有测试检测器,包括基于风格的检测器。但进一步发现,随着分析文档数量增加,人机文本分布再次可区分。整体表明,可靠检测需从单文档转向多文档分析。
原文摘要 · Abstract (English)
Despite considerable progress in the development of machine-text detectors, the ease with which machine-text can be manipulated to evade detection has led to suggestions that the problem is inherently intractable. In this work, we investigate the limits of such evasion strategies. We demonstrate that while current attacks, ranging from prompt engineering to detector-guided optimization can effectively degrade performance of standard detectors, they fail to erase the underlying stylistic "fingerprints" of machine text. We show that few-shot detectors that utilize the stylistic feature space are robust to these evasion attempts, reliably detecting samples even from models explicitly tuned to prevent detection. This raises the question: does style represent a universal defense against machine-detection attacks? We demonstrate that the answer is "no'' by introducing a novel paraphrasing approach that simultaneously optimizes for undetectability and adherence to specific human styles. We show that unlike prior methods, this attack effectively evades all considered detectors, including those that utilize writing style. However, we find that this evasion is not absolute: as the number of documents available for analysis grows, the human and machine distributions become distinguishable again. Overall, our findings suggest that reliable machine-text detection requires moving beyond single-document analysis to multi-document analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。