让机器人在离线学习中同时保持风格一致和高任务表现
Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment

- 用子轨迹标签统一定义行为风格,解决风格与奖励的冲突
- 新方法SCIQL在多个数据集上同时优于现有离线强化学习方法
- 适合需要稳定风格输出的智能体训练场景
我们研究基于显式风格监督(通过子轨迹标签函数)的离线强化学习,以生成风格可控的策略。在此设置下,由于分布偏移及风格与奖励间的固有矛盾,实现风格与高任务性能的对齐尤为困难。现有方法虽提出多种风格定义,但难以有效调和二者目标。为此,我们提出统一的行为风格定义,并构建实用框架。基于此,提出风格条件隐式Q学习(SCIQL),融合回溯重标注和价值学习等离线目标条件强化学习技术,结合新型门控优势加权回归机制,高效优化任务性能的同时保持风格对齐。实验表明,SCIQL在双目标上均显著优于先前离线方法。代码、数据集与可视化见:https://mathieu-petitbois.github.io/projects/sciql/。
原文摘要 · Abstract (English)
We study offline reinforcement learning of style-conditioned policies using explicit style supervision via subtrajectory labeling functions. In this setting, aligning style with high task performance is particularly challenging due to distribution shift and inherent conflicts between style and reward. Existing methods, despite introducing numerous definitions of style, often fail to reconcile these objectives effectively. To address these challenges, we propose a unified definition of behavior style and instantiate it into a practical framework. Building on this, we introduce Style-Conditioned Implicit Q-Learning (SCIQL), which leverages offline goal-conditioned RL techniques, such as hindsight relabeling and value learning, and combine it with a new Gated Advantage Weighted Regression mechanism to efficiently optimize task performance while preserving style alignment. Experiments demonstrate that SCIQL achieves superior performance on both objectives compared to prior offline methods. Code, datasets and visuals are available in: https://mathieu-petitbois.github.io/projects/sciql/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。