通过轨迹状态建模,提前发现多轮大模型代理的潜在安全风险。
TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling

- 基于观察模型隐藏层表示,学习每一步的潜在风险状态。
- 在多个基准上提升完整轨迹安全预测与主动风险识别能力。
- 无需逐步标注风险,适合长时序代理的安全训练与审计。
大模型代理在多轮工具使用和环境交互中运行,安全风险常出现在最终结果之前。事后审计难以及时发现风险。我们提出TRACES,一种基于表示的主动审计方法,通过观察模型的隐藏表示学习前缀级轨迹风险状态。TRACES从步骤表示中提取潜在机制特征,并建模其时间演化,判断部分轨迹是否正偏离至不安全行为。为避免高成本且模糊的步骤级风险标注,TRACES采用弱轨迹级监督训练,仍可生成密集的前缀级风险估计。在多个代理安全基准上,TRACES显著提升了全轨迹安全预测与主动风险辨别能力。分析还表明,这些风险状态可用于训练更安全的代理,凸显了主动审计在长时序代理安全中的广泛潜力。
原文摘要 · Abstract (English)
LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long before they surface in the final outcome. Reactive auditing is therefore insufficient: post-hoc diagnosis frequently misses the chance to flag risks while they are unfolding. We propose TRACES, a representation-based proactive auditor that learns prefix-level trajectory risk states from the hidden representations of an observer LLM. TRACES induces latent mechanism features from step representations and models their temporal evolution to estimate whether a partial trajectory is drifting toward unsafe behavior. To sidestep the cost and ambiguity of step-level risk annotation, TRACES is trained with weak trajectory-level supervision while still producing dense prefix-level risk estimates. Across multiple agent safety benchmarks, TRACES improves both full-trajectory safety prediction and proactive risk discrimination. Our analyses further suggest that these risk states can help train a safer agent, highlighting the broader potential of proactive auditing for long-horizon agent safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。