arXiv:2601.06663cs.AI2026-01被引 6

评测专业级AI代理的安全性,发现其在复杂任务中存在严重安全漏洞。

SafePro: Evaluating the Safety of Professional-Level AI Agents

  • 构建跨领域的高复杂度专业任务数据集,用于评估AI代理安全对齐。
  • 主流模型在专业任务中表现出明显安全判断不足与对齐缺陷。
  • 验证安全缓解策略有效,适合关注AI安全的开发者与研究者。

基于大语言模型的AI代理正从简单对话助手演变为能执行复杂专业任务的自主系统。尽管这些进展有望大幅提升生产力,却也带来了尚未充分探索的安全风险。现有安全评估多聚焦于日常辅助任务,难以捕捉专业场景中复杂决策过程及行为偏差带来的潜在后果。为此,我们提出 extbf{SafePro},一个全面的基准测试体系,用于评估执行专业活动的AI代理的安全对齐程度。SafePro包含多个领域、高复杂度的任务数据集,通过严谨的迭代创建与评审流程构建。对前沿AI模型的评估揭示了显著的安全漏洞,并发现了专业情境下的新型不安全行为。进一步分析表明,这些模型在执行复杂专业任务时,既缺乏足够的安全判断力,又表现出弱安全对齐。此外,我们探究了提升代理安全性的缓解策略,观察到令人鼓舞的改进效果。综合来看,研究凸显了为下一代专业级AI代理设计强健安全机制的紧迫性。

原文摘要 · Abstract (English)

Large language model-based agents are rapidly evolving from simple conversational assistants into autonomous systems capable of performing complex, professional-level tasks in various domains. While these advancements promise significant productivity gains, they also introduce critical safety risks that remain under-explored. Existing safety evaluations primarily focus on simple, daily assistance tasks, failing to capture the intricate decision-making processes and potential consequences of misaligned behaviors in professional settings. To address this gap, we introduce \textbf{SafePro}, a comprehensive benchmark designed to evaluate the safety alignment of AI agents performing professional activities. SafePro features a dataset of high-complexity tasks across diverse professional domains with safety risks, developed through a rigorous iterative creation and review process. Our evaluation of state-of-the-art AI models reveals significant safety vulnerabilities and uncovers new unsafe behaviors in professional contexts. We further show that these models exhibit both insufficient safety judgment and weak safety alignment when executing complex professional tasks. In addition, we investigate safety mitigation strategies for improving agent safety in these scenarios and observe encouraging improvements. Together, our findings highlight the urgent need for robust safety mechanisms tailored to the next generation of professional AI agents.

AI安全代理评估专业任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。