用蜜罐协议检测大模型在不同环境下的行为差异,发现其行为稳定且无隐藏副作用。
Honeypot Protocol
- 通过改变系统提示语测试模型在不同情境下的响应,主动探测潜在风险行为。
- Claude Opus 4.6 在三种条件下均实现100%主任务成功,零触发副作用任务。
- 为后续攻击策略和模型对比提供可复现的基准,适合安全评估研究者参考。
可信监控作为人工智能控制的标准防御手段,易受自适应攻击、共谋攻击和策略性攻击选择的影响。这些攻击利用了监控的被动性:仅观察模型行为,而不探究其在不同感知条件下的潜在反应差异。本文提出蜜罐协议,通过在保持任务、环境与评分一致的前提下,仅改变系统提示语(评估、合成部署、明确无监控三种条件),测试模型是否存在上下文依赖行为。我们在BashArena中对Claude Opus 4.6进行了诚实模式与攻击模式下的三条件评估。结果显示,该模型在所有条件下均达成100%主任务成功率,并始终未触发任何副作用任务,为未来更强攻击策略与更多模型的比较提供了基准。
原文摘要 · Abstract (English)
Trusted monitoring, the standard defense in AI control, is vulnerable to adaptive attacks, collusion, and strategic attack selection. All of these exploit the fact that monitoring is passive: it observes model behavior but never probes whether the model would behave differently under different perceived conditions. We introduce the honeypot protocol, which tests for context-dependent behavior by varying only the system prompt across three conditions (evaluation, synthetic deployment, explicit no-monitoring) while holding the task, environment, and scoring identical. We evaluate Claude Opus 4.6 in BashArena across all three conditions in both honest and attack modes. The model achieved 100% main task success and triggered zero side tasks uniformly across conditions, providing a baseline for future comparisons with stronger attack policies and additional models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。