arXiv:2602.03100cs.AI2026-02被引 7

构建真实部署下的智能体安全评估框架,发现主流模型存在显著风险。

Risky-Bench: Probing Agentic Safety Risks under Real-World Deployment

  • 基于通用安全原则设计情境感知的评估标准
  • 在真实任务中暴露先进智能体的安全隐患
  • 可扩展至多场景,适合安全研究者与开发者使用

大型语言模型作为智能体在真实环境中部署时,带来超越语言伤害的安全风险。现有评估方法依赖特定场景的风险任务,覆盖范围有限,且难以在复杂、长周期交互中评估安全行为。为此,我们提出Risky-Bench框架,基于领域无关的安全原则,生成情境感知的安全准则,系统评估不同威胁假设下真实任务执行中的安全风险。应用于生活辅助智能体场景时,Risky-Bench揭示了当前顶尖智能体在真实运行条件下的重大安全隐患。该框架结构清晰,不仅适用于生活辅助场景,还可拓展至其他部署环境,为智能体安全评估提供可复用的方法论。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed as agents that operate in real-world environments, introducing safety risks beyond linguistic harm. Existing agent safety evaluations rely on risk-oriented tasks tailored to specific agent settings, resulting in limited coverage of safety risk space and failing to assess agent safety behavior during long-horizon, interactive task execution in complex real-world deployments. Moreover, their specialization to particular agent settings limits adaptability across diverse agent configurations. To address these limitations, we propose Risky-Bench, a framework that enables systematic agent safety evaluation grounded in real-world deployment. Risky-Bench organizes evaluation around domain-agnostic safety principles to derive context-aware safety rubrics that delineate safety space, and systematically evaluates safety risks across this space through realistic task execution under varying threat assumptions. When applied to life-assist agent settings, Risky-Bench uncovers substantial safety risks in state-of-the-art agents under realistic execution conditions. Moreover, as a well-structured evaluation pipeline, Risky-Bench is not confined to life-assist scenarios and can be adapted to other deployment settings to construct environment-specific safety evaluations, providing an extensible methodology for agent safety assessment.

智能体安全评估框架真实部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。