用四维评分框架提升智能体工具使用安全性
RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

- 将智能体行为拆解为工具、论证、回复和帮助性四维度评分
- 相比基线方法,安全提升且工具幻觉减少37%
- 适合需高安全性的工具调用类智能体开发
大型语言模型演变为具备工具调用能力的智能体,带来了与真实世界执行相关的新型安全挑战,而非仅限于文本生成。现有对齐方法常依赖粗粒度拒绝信号或静态监督,难以在多样化的智能体风险下平衡安全与有用性。本文提出RUBAS,一种基于评分标准的强化学习框架,用于智能体安全。RUBAS将智能体行为分解为四个维度:工具使用安全、论证安全、响应安全与助人性。这些结构化评分标准提供完整智能体轨迹上的细粒度、可解释奖励,使强化学习能够优化安全的工具使用,同时保持任务完成度。在多个智能体安全基准和模型上的广泛实验表明,RUBAS在安全性上优于标准对齐基线,显著减少工具相关幻觉,并维持有竞争力的实用性。结果表明,多维度评分奖励能有效引导在关键工具使用场景中对齐大型语言模型智能体。
原文摘要 · Abstract (English)
The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text generation. Existing alignment methods often rely on coarse refusal signals or static supervision, making it difficult to balance safety with useful tool execution across diverse agentic risks. We introduce RUBAS, a rubric-based reinforcement learning framework for agent safety. RUBAS decomposes agent behavior into four dimensions: tool-use safety, argument safety, response safety, and helpfulness. These structured rubrics provide fine-grained and interpretable rewards over complete agent trajectories, enabling reinforcement learning to optimize safe tool use while preserving task completion. Extensive experiments across multiple agent safety benchmarks and models show that RUBAS improves safety over standard alignment baselines, reduces tool-grounded hallucinations, and maintains competitive utility. Our results suggest that multi-dimensional rubric rewards provide an effective training signal for aligning LLM agents in safety-critical tool-use settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。