首个评估机器文本伪装攻击的基准,揭示三重权衡。
TH-Bench: Evaluating Evading Attacks via Humanizing AI Text on Machine-Generated Text Detectors
- 构建多维度评估框架,测试攻击在逃逸性、质量与开销上的表现。
- 6种先进攻击在13个检测器上测试,发现无一在三方面全优。
- 揭示攻击的三重权衡,提出优化方向供后续研究参考。
随着大语言模型(LLMs)的发展,机器生成文本(MGT)愈发流畅、高质量且信息丰富。现有广泛使用的MGT检测器旨在识别此类文本以防止抄袭和虚假信息传播。然而,攻击者尝试通过轻微修改使MGT更像人类写作(称为逃逸攻击),从而规避检测。当前攻击缺乏统一的评估框架,因实验设置、模型架构和数据集各异而难以比较。为此,我们提出文本人性化基准(TH-Bench),首个全面评估逃逸攻击的基准。TH-Bench从逃逸有效性、文本质量与计算开销三个维度进行评估。我们对6种顶尖攻击方法,在13个检测器上,跨6个数据集、19个领域、由11种主流LLM生成的文本进行了广泛实验。结果表明,无一攻击在三方面均占优。深入分析揭示了各类攻击的优劣,并发现三者间存在显著权衡。基于此,我们提出两项优化洞察,初步实验验证其正确性与有效性,为未来研究提供新方向。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) advance, Machine-Generated Texts (MGTs) have become increasingly fluent, high-quality, and informative. Existing wide-range MGT detectors are designed to identify MGTs to prevent the spread of plagiarism and misinformation. However, adversaries attempt to humanize MGTs to evade detection (named evading attacks), which requires only minor modifications to bypass MGT detectors. Unfortunately, existing attacks generally lack a unified and comprehensive evaluation framework, as they are assessed using different experimental settings, model architectures, and datasets. To fill this gap, we introduce the Text-Humanization Benchmark (TH-Bench), the first comprehensive benchmark to evaluate evading attacks against MGT detectors. TH-Bench evaluates attacks across three key dimensions: evading effectiveness, text quality, and computational overhead. Our extensive experiments evaluate 6 state-of-the-art attacks against 13 MGT detectors across 6 datasets, spanning 19 domains and generated by 11 widely used LLMs. Our findings reveal that no single evading attack excels across all three dimensions. Through in-depth analysis, we highlight the strengths and limitations of different attacks. More importantly, we identify a trade-off among three dimensions and propose two optimization insights. Through preliminary experiments, we validate their correctness and effectiveness, offering potential directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。