机器文本检测器本质是会员推断攻击,二者可相互迁移。
Machine Text Detectors are Membership Inference Attacks
- 发现两类任务共享最优度量标准,理论统一
- 实验证明跨任务性能相关性高达0.7,检测器在两任务中表现最强
- 提出MINT评估套件,支持双任务公平比较
尽管会员推断攻击(MIAs)与机器生成文本检测的目标不同,但其方法常依赖语言模型概率分布的相似信号,且两个方向长期独立研究,可能忽略彼此的强方法与洞见。本文从理论上和实验上证明了两类任务之间的可迁移性:我们证明两者达到渐近最优性能的度量完全一致,并在此基础上统一现有方法。假设方法对这一最优度量的逼近程度直接决定其迁移能力。大规模实验显示跨任务性能具有极强秩相关性(ρ≈0.7)。特别地,一个机器文本检测器在两项任务中均取得最佳表现,凸显迁移的实际价值。为促进跨任务发展与公平评估,我们构建MINT——一个统一评估套件,集成来自两类任务的15个近期方法。
原文摘要 · Abstract (English)
Although membership inference attacks (MIAs) and machine-generated text detection target different goals, their methods often exploit similar signals based on a language model's probability distribution, and the two tasks have been studied independently. This can result in conclusions that overlook stronger methods and valuable insights from the other task. In this work, we theoretically and empirically demonstrate the transferability, i.e., how well a method originally developed for one task performs on the other, between MIAs and machine text detection. We prove that the metric achieving asymptotically optimal performance is identical for both tasks. We unify existing methods under this optimal metric and hypothesize that the accuracy with which a method approximates this metric is directly correlated with its transferability. Our large-scale empirical experiments demonstrate very strong rank correlation ($ρ\approx 0.7$) in cross-task performance. Notably, we also find that a machine text detector achieves the strongest performance among evaluated methods on both tasks, demonstrating the practical impact of transferability. To facilitate cross-task development and fair evaluation, we introduce MINT, a unified evaluation suite for MIAs and machine-generated text detection, implementing 15 recent methods from both tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。