让大模型在推理中主动识别潜在危害,避免无意助人作恶。
AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models
- 用分步奖励模型动态评估每一步推理的安全性与逻辑一致性。
- 实验显示该方法显著提升输出逻辑完整性和对潜在风险的敏感度。
- 适合需要高安全性的AI应用,如医疗、金融等决策场景。
当前的大语言模型面临基于可用性(affordance)的安全风险——即在未察觉逻辑后果的情况下,输出可能无意中促成有害行为。传统安全方案如基于标量结果的奖励模型、参数调优或启发式解码策略,缺乏足够粒度和主动性,难以在细微但关键的推理步骤中可靠地检测并干预。为填补这一根本性空白,我们提出AURA,一种以过程奖励模型(Process Reward Models, PRMs)为核心的多层框架,实现对逻辑连贯性和安全意识的全流程、细粒度评估。该框架融合内省式自我批判、精细的PRM评估与自适应安全感知解码,动态且主动地引导模型走向更安全的推理路径。实证表明,该方法显著优于现有方法,在提升输出逻辑完整性与对可用性敏感安全性的表现上均取得突破。本研究标志着迈向更安全、更负责任、更具情境意识的人工智能的重要一步,为对齐敏感型应用设立了新基准。
原文摘要 · Abstract (English)
Present day LLMs face the challenge of managing affordance-based safety risks-situations where outputs inadvertently facilitate harmful actions due to overlooked logical implications. Traditional safety solutions, such as scalar outcome-based reward models, parameter tuning, or heuristic decoding strategies, lack the granularity and proactive nature needed to reliably detect and intervene during subtle yet crucial reasoning steps. Addressing this fundamental gap, we introduce AURA, an innovative, multi-layered framework centered around Process Reward Models (PRMs), providing comprehensive, step level evaluations across logical coherence and safety-awareness. Our framework seamlessly combines introspective self-critique, fine-grained PRM assessments, and adaptive safety-aware decoding to dynamically and proactively guide models toward safer reasoning trajectories. Empirical evidence clearly demonstrates that this approach significantly surpasses existing methods, significantly improving the logical integrity and affordance-sensitive safety of model outputs. This research represents a pivotal step toward safer, more responsible, and contextually aware AI, setting a new benchmark for alignment-sensitive applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。