自动生成1072个安全场景,评估大模型代理执行中的真实风险。
ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

- 基于五大风险维度自动生成可测量的评估场景
- 12个代理平均行为安全率仅47.1%,部分超70%
- 适合关注大模型安全评估与工程落地的研究者
大型语言模型正从简单的文本交互系统演变为具备记忆、工具调用和环境交互能力的智能体。随着其能力与自主性的增强,面临的安全风险也日益复杂。现有评估多依赖人工编写场景、静态提示或最终输出判断,难以捕捉任务执行过程中的多样化风险。本文提出ForesightSafety-SAGE,一个面向大模型智能体的全自动场景生成与安全评估框架。基于五个风险维度,将现实任务执行中的抽象多样风险转化为1072个可测量的评估场景。通过自动化评估流程,在两种权限背景下对12个大模型智能体进行了评估。结果表明,当前智能体在任务执行中仍面临显著的行为安全风险,平均安全率(ASR)为47.1%,部分模型超过70%。研究凸显了可执行、过程级评估对于理解并提升大模型智能体安全性的关键作用。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks. As their capabilities and autonomy expand, the safety risks they face also become more diverse. Existing evaluations often rely on manually written scenarios, static prompts, or final-output judgments, making it difficult to capture the diverse risks that agents may face during task execution. We introduce ForesightSafety-SAGE, a fully automated scenario generation and safety evaluation framework for LLM agents. Based on five risk dimensions,we instantiae abstract and diverse safety risks in real-world task execution into 1,072 measurable evaluation scenarios. Using the automated evaluation pipeline, 12 LLM agents are evaluated under two authority contexts. The results show that current agents still face substantial behavioral safety risks during task execution, with an average ASR of 47.1% and several models exceeding 70%. These findings demonstrate the importance of executable, process-level evaluation for understanding and improving LLM agent safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。