分离奖励与策略,让AI对齐更可查可改可复用
Interactionless Inverse Reinforcement Learning: A Data-Centric Framework for Durable Alignment
- 不修改策略参数,直接学习独立可读的奖励模型
- 通过人工反馈循环优化,实现对齐成果持续迭代
- 适合需要长期维护、可审计AI系统的团队使用
AI对齐日益重要,但现有方法常通过直接修改策略参数来学习安全行为,导致规范约束与策略混杂,形成难以理解、不可编辑、无法复用的对齐结果,我们称之为对齐浪费。本文提出无交互逆强化学习框架,将可检查、可编辑、可复用的奖励模型从策略优化中分离出来。进一步引入对齐飞轮机制,通过自动化评估与人工反馈循环,实现对齐成果的持续审计、修补与强化。该方法将对齐从一次性训练开销转变为持久、可验证的工程资产。
原文摘要 · Abstract (English)
AI alignment is growing in importance, yet many current approaches learn safety behavior by directly modifying policy parameters, entangling normative constraints with the underlying policy. This often yields opaque, difficult-to-edit alignment artifacts and reduces their reuse across models or deployments, a failure mode we term Alignment Waste. We propose Interactionless Inverse Reinforcement Learning, a framework for learning inspectable, editable, and reusable reward artifacts separately from policy optimization. We further introduce the Alignment Flywheel, a human-in-the-loop lifecycle for iteratively auditing, patching, and hardening these artifacts through automated evaluation and refinement. Together, these ideas recast alignment from a disposable training expense into a durable, verifiable engineering asset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。