用真实操作轨迹训练安全防护模型,应对动态演化的计算机使用威胁。
BraveGuard: From Open-World Threats to Safer Computer-Use Agents

- 从真实威胁中自动挖掘攻击模式,生成可执行任务用于训练
- 在AgentHazard上将检测准确率从38.79%提升至82.38%
- 适合关注智能代理安全的开发者与研究者
计算机使用代理将语言模型从文本生成拓展到与文件、终端、浏览器及外部工具的持续交互。这一转变带来了难以通过孤立提示或最终响应发现的安全风险,因为危害往往在多步执行轨迹中逐步显现,单个动作看似无害。我们提出BraveGuard,一种基于开放世界威胁信号和真实代理轨迹自演化防御框架。BraveGuard通过挖掘近期研究源识别新兴风险与攻击模式,将其转化为可执行的计算机使用任务,收集代理轨迹,并生成轨迹级监督信号用于守护模型训练。当新威胁或验证失败出现时,该流程可重复运行,形成动态适应的防御闭环,而非静态基准驱动的训练过程。我们通过训练多个守护模型骨干(包括Qwen3-Guard和Llama-Guard变体)并评估其在轨迹级代理安全基准上的表现,结果表明BraveGuard在多种计算机使用轨迹中持续提升安全检测能力。在AgentHazard上,平均守护模型准确率从38.79%提升至82.38%。这表明,基于开放世界威胁发现与真实代理执行的守护监督,可超越固定分类体系与合成提示数据,在安全监控方面实现显著改进。BraveGuard为应对不断演变的真实世界风险,提供了可扩展的自适应防御路径。
原文摘要 · Abstract (English)
Computer-use agents extend language models from text generation to sustained interaction with files, terminals, browsers, and external tools. This shift creates safety risks that are difficult to detect from isolated prompts or final responses, because harm often emerges only through multi-step execution traces whose individual actions appear locally benign. We introduce BraveGuard, a self-evolving defense framework for training guard models from open-world threat signals and realistic agent trajectories. BraveGuard mines recent research sources to identify emerging risks and attack patterns, instantiates them as executable computer-use tasks, collects agent rollouts, and derives trajectory-level supervision for guard model training. As new threats and validation failures appear, the pipeline can be repeated, yielding an adaptive defense loop rather than a static, benchmark-driven training process. We instantiate BraveGuard by training multiple guard backbones, including Qwen3-Guard and Llama-Guard variants, and evaluate the resulting guards on trajectory-level agent-safety benchmarks. BraveGuard consistently improves safety detection across computer-use trajectories. On AgentHazard, it substantially improves detection accuracy over off-the-shelf guard models, with accuracy increasing from 38.79% to 82.38% under the averaged guard-model setting. These results show that guard supervision grounded in open-world threat discovery and realistic agent execution can improve safety monitoring beyond fixed taxonomies and synthetic prompt-level data. BraveGuard offers a scalable path toward adaptive defenses for computer-use agents facing evolving real-world risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。