arXiv:2605.24216cs.LGcs.AI2026-05

用心理模型分析大模型代理,提前发现隐藏恶意行为

Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning

论文配图:Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning
图 1 · 摘自论文原文
  • 基于心理模型推理,追踪代理的信念与意图
  • 在两个基准测试中精度与召回率均优于现有方法
  • 适合关注大模型安全监控的研究者与工程师

监控自主大型语言模型(LLM)代理的隐蔽恶意行为极具挑战,因其攻击模式具有延迟性、依赖上下文和长时程特征。代理可能在表面看似正常执行任务的同时,暗中追求隐藏目标,即使拥有完整行为轨迹也难以察觉。现有监控方法多依赖结构设计或集成聚合,但独立处理每条轨迹,无法从过往经验中学习。此外,标准推理方法仅解释观察到的行为,未显式推理代理的信念、意图与目标一致性,难以区分良性执行与隐蔽偏离。本文提出 extbf{Agent-ToM},一种基于心理模型(ToM)推理的可学习监控框架,用于自主代理的安全分析。该框架通过推断信念、带置信度的意图假设、预期动作及与任务一致行为基线的偏差,进行全轨迹结构化分析。推理时采用 extit{Reason-Verify-Refine} 流水线构建并验证监控决策;训练时将批判信号提炼为持久的 extit{语义护栏记忆},实现跨任务的信念与意图条件约束复用。我们在对抗性代理监控基准(SHADE-Arena 与 CUA-SHADE-Arena)上评估该方法,结果表明,Agent-ToM 在精度-召回平衡上表现优异,超越当前最优监控基线(包括集成方法),且仅需两次调用推理流程。实验验证了在监控层学习结合结构化心理模型推理与验证,是保障自主LLM代理安全的有效且可部署方案。

原文摘要 · Abstract (English)

Monitoring autonomous large language model (LLM) agents for covert malicious behavior is challenging due to delayed, context-dependent, and long-horizon attack patterns. Agents may pursue hidden objectives while maintaining superficially benign behavior, making detection difficult even with full trajectory access. Prior monitoring approaches improve scaffolding or ensemble aggregation, but treat each trajectory independently and do not learn from prior monitoring experience. Moreover, standard reasoning methods explain observed behavior without explicitly reasoning about agent beliefs, intentions, and goal alignment required to distinguish benign task execution from covert deviation. We propose \textbf{Agent-ToM}, a learning-to-monitor framework grounded in Theory-of-Mind (ToM) reasoning for security analysis of autonomous agents. Agent-ToM performs structured full-trajectory analysis by inferring beliefs, intent hypotheses with calibrated confidence, expected actions, and deviations from task-consistent behavioral baselines. At inference time, it employs a \textit{Reason-Verify-Refine} pipeline to construct and validate monitoring decisions. At training time, Agent-ToM distills critique signals into a persistent \textit{semantic guardrail memory}, enabling reusable belief- and intent-conditioned constraints across episodes. We evaluate Agent-ToM on adversarial agent monitoring benchmarks (SHADE-Arena and CUA-SHADE-Arena). Agent-ToM achieves strong precision-recall balance and outperforms state-of-the-art monitoring baselines, including ensemble methods, while using a coherent two-call reasoning pipeline. These results demonstrate that learning at the monitoring layer, combined with structured ToM reasoning and verification, provides an effective and deployable foundation for securing autonomous LLM agents.

大模型安全心理模型监控框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。