构建欺骗性安全测试集,提升大模型对隐性风险的判断能力
Enhancing Agent Safety Judgment: Controlled Benchmark Rewriting and Analogical Reasoning for Deceptive Out-of-Distribution Scenarios

- 通过重写100条危险轨迹生成300个更具迷惑性的测试用例
- 新测试集使前沿模型安全判断性能显著下降,隐性风险仍难识别
- 引入类比推理机制,在不重新训练情况下增强判断鲁棒性
基于大语言模型的工具使用智能体正广泛部署于网络、应用、操作系统及交易环境。然而现有安全评估基准仍侧重显性风险,可能高估模型对欺骗性或模糊行为的判断能力。为此,我们提出ROME(红队驱动的多智能体演化)——一种受控的基准构建流程,将已知的100条危险轨迹改写为更具有欺骗性的评估实例,同时保持其原始风险标签。该方法生成了300个涵盖上下文模糊、隐性风险与捷径决策的挑战样本。实验表明,这些挑战集显著降低安全判断性能,即使在最新前沿模型中,隐性风险案例依然难以处理。我们进一步研究ARISE(类比推理用于推理时安全增强),一种基于检索的推理时增强机制,从外部类比库中检索类似ReAct风格的安全轨迹,并注入为结构化推理示例。ARISE可在不重新训练的前提下提升判断质量,但应视为任务特定的鲁棒性增强,而非独立的安全保障。整体而言,ROME与ARISE为在欺骗性分布偏移下压力测试和改进智能体安全判断提供了实用工具。
原文摘要 · Abstract (English)
Tool-using agent systems powered by large language models (LLMs) are increasingly deployed across web, app, operating-system, and transactional environments. Yet existing safety benchmarks still emphasize explicit risks, potentially overstating a model's ability to judge deceptive or ambiguous trajectories. To address this gap, we introduce ROME (Red-team Orchestrated Multi-agent Evolution), a controlled benchmark-construction pipeline that rewrites known unsafe trajectories into more deceptive evaluation instances while preserving their underlying risk labels. Starting from 100 unsafe source trajectories, ROME produces 300 challenge instances spanning contextual ambiguity, implicit risks, and shortcut decision-making. Experiments show that these challenge sets substantially degrade safety-judgment performance, with hidden-risk cases remaining particularly non-trivial even for recent frontier models. We further study ARISE (Analogical Reasoning for Inference-time Safety Enhancement), a retrieval-guided inference-time enhancement that retrieves ReAct-style analogical safety trajectories from an external analogical base and injects them as structured reasoning exemplars. ARISE improves judgment quality without retraining, but is best viewed as a task-specific robustness enhancement rather than a standalone safety guarantee. Together, ROME and ARISE provide practical tools for stress-testing and improving agent safety judgment under deceptive distribution shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。