arXiv:2512.17259cs.MAcs.AI2025-12被引 3

让大模型代理行为可验证,快速发现并纠正偏差。

Verifiability-First Agents: Provable Observability and Lightweight Audit Agents for Controlling Autonomous LLM Systems

  • 用密码学与符号方法实时证明代理动作,确保可追溯。
  • 部署轻量审计代理,持续比对意图与实际行为,检测偏差。
  • 高风险操作需通过挑战-响应验证,适合安全敏感场景。

随着基于大模型的代理日益自主和多模态,确保其可控、可审计且忠实于部署者意图变得至关重要。先前基准测试显示,代理性格和工具访问权显著影响行为偏离。本文提出可信优先架构:(1)利用密码学与符号方法在运行时对代理行为进行实时证明;(2)嵌入轻量级审计代理,通过受限推理持续验证意图与行为的一致性;(3)对高风险操作强制执行挑战-响应证明协议。我们引入OPERA(可观测性、可证明执行、红队测试、证明)基准套件与评估流程,用于衡量(i)行为偏离的可检测性,(ii)隐蔽策略下的检测时间,以及(iii)验证机制对抗恶意提示与角色注入的鲁棒性。本方法将评估重点从偏离发生的概率转向偏差被快速可靠检测与修复的能力。

原文摘要 · Abstract (English)

As LLM-based agents grow more autonomous and multi-modal, ensuring they remain controllable, auditable, and faithful to deployer intent becomes critical. Prior benchmarks measured the propensity for misaligned behavior and showed that agent personalities and tool access significantly influence misalignment. Building on these insights, we propose a Verifiability-First architecture that (1) integrates run-time attestations of agent actions using cryptographic and symbolic methods, (2) embeds lightweight Audit Agents that continuously verify intent versus behavior using constrained reasoning, and (3) enforces challenge-response attestation protocols for high-risk operations. We introduce OPERA (Observability, Provable Execution, Red-team, Attestation), a benchmark suite and evaluation protocol designed to measure (i) detectability of misalignment, (ii) time to detection under stealthy strategies, and (iii) resilience of verifiability mechanisms to adversarial prompt and persona injection. Our approach shifts the evaluation focus from how likely misalignment is to how quickly and reliably misalignment can be detected and remediated.

大模型代理可验证性审计机制安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。