arXiv:2603.00381cs.CRcs.AI2026-03

通过可验证准入机制,防止语言模型代理隐藏协同信号。

Verifier-Bound Communication for LLM Agents: Certified Bounds on Covert Signaling

  • 生成与审核分离,用小规模验证器检查消息合规性
  • 实测显示隐蔽信道泄露极低,对抗攻击也未突破安全阈值
  • 适合关注大模型安全协作的系统设计者与研究者

合谋的语言模型代理可在表面合规的文本中隐藏协调信息。本文提出CLBC协议,将生成与准入分离:仅当小型验证器在固定谓词Π下接受一个证明绑定的包裹时,消息才被纳入对话状态。该谓词绑定策略哈希、公开随机调度、对话链式结构、潜在模式约束、规范元数据/工具字段及确定性拒绝码。我们推导出对话泄露上限,由潜在泄露加显式残余通道构成;给出自适应组合保证,并在策略合法选项仍可选时确立语义下界。实证结果显示:整体评估满足所有预设阈值;严格车道解码优势为0.0000,互信息代理值为0.0636;自适应合谋压力测试低于攻击者阈值;基线对比显示默认拒绝语义与仅审计控制间存在显著差距。进一步量化操作权衡:全证明模式中位数回合延迟为27.53秒(p95 28.08秒),而采样证明将非证明回合延迟降至0.327毫秒。核心发现是:仅靠瓶颈无法保障安全,安全声明依赖于在线、确定且故障封闭的可验证准入语义。

原文摘要 · Abstract (English)

Colluding language-model agents can hide coordination in messages that remain policy-compliant at the surface level. We present CLBC, a protocol where generation and admission are separated: a message is admitted to transcript state only if a small verifier accepts a proof-bound envelope under a pinned predicate $Π$. The predicate binds policy hash, public randomness schedule, transcript chaining, latent schema constraints, canonical metadata/tool fields, and deterministic rejection codes. We show how this protocol yields an upper bound on transcript leakage in terms of latent leakage plus explicit residual channels, derive adaptive composition guarantees, and state a semantic lower bound when policy-valid alternatives remain choosable. We report extensive empirically grounded evidence: aggregate evaluation satisfies all prespecified thresholds; strict lane decoder advantage is bounded at 0.0000 with MI proxy 0.0636; adaptive-colluder stress tests remain below attacker thresholds; and baseline separation shows large gaps between reject-by-default semantics and audit-only controls. We further quantify operational tradeoffs. Strict full-proof mode has median turn latency 27.53s (p95 28.08s), while sampled proving reduces non-proved-turn latency to 0.327ms. The central finding is that bottlenecks alone are insufficient: security claims depend on verifiable admission semantics that are online, deterministic, and fail-closed.

大模型安全可信验证协同防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。