arXiv:2507.15907cs.LGcs.AI2025-07

提出新型对抗框架,让AI难以被发现且保持高质量输出。

Dual Turing Test: A Framework for Detecting and Mitigating Undetectable AI

  • 以人类识别AI为目标,构建反向图灵测试框架
  • 结合质量约束与最坏情况保障,实现隐蔽性与流畅性的平衡
  • 适合关注AI安全与对齐的科研人员参考

本文提出一个统一框架,将三种方向融合:(1) 反向图灵测试,即人类裁判目标是识别AI而非判断其欺骗性;(2) 带有显式质量约束和最坏情况保证的对抗分类游戏;(3) 使用不可检测性检测器和质量相关组件的强化学习对齐流程。回顾从逆向到元图灵测试的历史先例,强调结合质量阈值、分阶段难度和极小极大边界的新颖性。形式化定义在N轮独立测试中,从提示空间Q中抽取新提示,引入质量函数Q、参数tau和delta,将交互建模为在对手可行策略集M上的双人零和博弈。进一步将其映射到类似RL-HF的对齐循环中,由不可检测性检测器D对隐蔽输出给予负奖励,由质量代理维持语言流畅性。详细解释各组件符号含义、序列内最小化意义、分阶段测试及迭代对抗训练,并提出若干立即行动建议。

原文摘要 · Abstract (English)

In this short note, we propose a unified framework that bridges three areas: (1) a flipped perspective on the Turing Test, the "dual Turing test", in which a human judge's goal is to identify an AI rather than reward a machine for deception; (2) a formal adversarial classification game with explicit quality constraints and worst-case guarantees; and (3) a reinforcement learning (RL) alignment pipeline that uses an undetectability detector and a set of quality related components in its reward model. We review historical precedents, from inverted and meta-Turing variants to modern supervised reverse-Turing classifiers, and highlight the novelty of combining quality thresholds, phased difficulty levels, and minimax bounds. We then formalize the dual test: define the judge's task over N independent rounds with fresh prompts drawn from a prompt space Q, introduce a quality function Q and parameters tau and delta, and cast the interaction as a two-player zero-sum game over the adversary's feasible strategy set M. Next, we map this minimax game onto an RL-HF style alignment loop, in which an undetectability detector D provides negative reward for stealthy outputs, balanced by a quality proxy that preserves fluency. Throughout, we include detailed explanations of each component notation, the meaning of inner minimization over sequences, phased tests, and iterative adversarial training and conclude with a suggestion for a couple of immediate actions.

AI安全对抗检测强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。