arXiv:2604.13954cs.LGcs.AI2026-04被引 6

提出新基准HINTBench,评估智能体在无外部攻击下的内在风险轨迹。

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

论文配图:HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark
图 1 · 摘自论文原文
  • 构建629条轨迹的基准,标注5类内在风险约束
  • 强LLM在轨迹级检测表现好,但步骤级定位准确率低于35%
  • 适合关注长时程智能体安全与失效诊断的研究者

现有智能体安全性评估主要关注外部诱导风险,但智能体在正常条件下仍可能进入危险轨迹。本文从内在风险视角研究这一被忽视的问题:内在故障潜伏、在长时程执行中传播,最终导致高后果结果。为此,我们提出非攻击性内在风险审计方法,并构建了包含629条智能体轨迹(其中523条具风险,106条安全;平均33步)的基准HINTBench,支持三项任务:风险检测、风险步骤定位与内在故障类型识别。其标注基于统一的五约束分类体系。实验表明存在显著能力差距:强语言模型在轨迹级风险检测表现良好,但在步骤级定位上严格F1低于35%,细粒度故障诊断更难。现有防护模型在此场景下迁移效果差。这些发现确立了内在风险审计作为智能体安全领域的一项开放挑战。

原文摘要 · Abstract (English)

Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions. We study this complementary but underexplored setting through the lens of \emph{intrinsic} risk, where intrinsic failures remain latent, propagate across long-horizon execution, and eventually lead to high-consequence outcomes. To evaluate this setting, we introduce \emph{non-attack intrinsic risk auditing} and present \textbf{HINTBench}, a benchmark of 629 agent trajectories (523 risky, 106 safe; 33 steps on average) supporting three tasks: risk detection, risk-step localization, and intrinsic failure-type identification. Its annotations are organized under a unified five-constraint taxonomy. Experiments reveal a substantial capability gap: strong LLMs perform well on trajectory-level risk detection, but their performance drops to below 35 Strict-F1 on risk-step localization, while fine-grained failure diagnosis proves even harder. Existing guard models transfer poorly to this setting. These findings establish intrinsic risk auditing as an open challenge for agent safety.

智能体安全风险评估长时程轨迹基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。