arXiv:2607.07859cs.AIcs.HC2026-07

用反馈修正信号提升离线模仿学习的对齐效果,减少98%行为偏差。

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

论文配图:Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning
图 1 · 摘自论文原文
  • 利用评估反馈作为纠正信号,改进模仿学习策略对齐
  • 在多种算法上实现最高98%的不对齐降低
  • 适用于数据稀疏、噪声示范的现实场景

强化学习研究日益关注对齐问题,确保智能体行为符合人类价值观。尽管人类示范和反馈对对齐至关重要,现有方法多采用分阶段流水线处理,适用于语言生成的上下文带宽框架。然而,很少有工作探索如何将这些互补输入作为更丰富、互联的信号,在单阶段离线训练中用于完全序列决策环境。本文提出反馈操控正则化(FMR),一种与算法无关的方法,利用评估反馈作为校正信号,改善模仿学习策略的对齐性。我们改造了Safety Gymnasium环境,构建了对齐评估的严谨测试平台,在多种模仿学习算法上展示出更强的适配能力,最多实现98%的不对齐减少。FMR在数据有限情况下仍具鲁棒性,即使在少量对齐示范和无信息噪声示范下也能有效学习。

原文摘要 · Abstract (English)

Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While human demonstrations and feedback have proven crucial for alignment, existing approaches predominantly combine these signals using multi-stage pipelines designed for the contextual bandit framing of language generation. Yet little work explores how these complementary inputs can serve as a richer, interconnected signal for single-stage offline training in fully sequential decision-making environments. We propose Feedback Manipulation Regularization (FMR), an algorithm-agnostic method that harnesses evaluative feedback as a corrective signal to improve the alignment of imitation learning policies. We adapt Safety Gymnasium environments to be a principled testbed for alignment evaluation, demonstrating improved aptitude and up to a 98\% reduction in misalignment across a range of imitation learning algorithms. FMR remains robust in limited data regimes, even when learning from scarce aligned and uninformative noisy demonstrations.

模仿学习对齐离线训练反馈正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。