arXiv:2504.02725cs.CL2025-04ACL被引 3

通过提前推理提升大模型安全对齐,兼顾安全与效率。

SAFER: Advancing Safety Alignment via Efficient Ex-Ante Reasoning

  • 引入结构化提前推理机制,分三步评估风险
  • 在多个开源模型上实现安全性能显著提升
  • 适合关注大模型安全落地的研究者与工程师

大型语言模型(LLMs)的快速发展推动了通用人工智能的进步,但其生成有害内容的潜在风险带来了严峻的安全挑战。现有对齐方法难以覆盖多样化的安全场景,且易受对抗攻击影响。本文提出 SAFER 框架,通过高效事前推理实现安全对齐。该方法通过初始评估、规则验证和路径校准实现结构化的事前推理,并嵌入预设安全规则,提供透明可验证的安全判断。具体包含两个训练阶段:(1) 使用合成轨迹进行监督微调,教授多阶段事前推理;(2) 逐步推理偏好优化,联合提升安全性、实用性与效率。在多个开源 LLM 上的实验表明,SAFER 显著提升了安全表现,同时保持有用性和响应效率。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have accelerated progress toward artificial general intelligence, yet their potential to generate harmful content poses critical safety challenges. Existing alignment methods often struggle to cover diverse safety scenarios and remain vulnerable to adversarial attacks. In this work, we propose SAFER, a framework for Safety Alignment via eFficient Ex-Ante Reasoning. Our approach instantiates structured Ex-Ante reasoning through initial assessment, rule verification, and path calibration, and embeds predefined safety rules to provide transparent and verifiable safety judgments. Specifically, our approach consists of two training stages: (1) supervised fine-tuning with synthetic traces to teach the multi-stage Ex-Ante reasoning, and (2) step-level reasoning preference optimization to jointly enhance safety, utility, and efficiency. Experiments on multiple open-source LLMs demonstrate that SAFER significantly enhances safety performance while maintaining helpfulness and response efficiency.

大模型安全推理对齐事前推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。