arXiv:2606.23280cs.RO2026-06

用因果模型实现零样本奖励设计,让机器人自动学会新技能

Causal Reward World Models: Zero-shot Reward Design for Automated Skill Generation

论文配图:Causal Reward World Models: Zero-shot Reward Design for Automated Skill Generation
图 1 · 摘自论文原文
  • 构建因果奖励世界模型,从多任务数据中学习奖励组件与目标变量的因果关系
  • 零样本生成可执行奖励函数,在复杂连续控制任务上性能达顶尖水平
  • 适合需要快速部署新机器人技能的研究者,提升自动化程度和泛化能力

自动化奖励设计(ARD)旨在用语言驱动的奖励函数生成替代强化学习中的手动奖励工程。然而,现有基于大语言模型(LLMs)的方法仍依赖相关性,需通过环境反馈迭代优化每个任务的奖励假设,不仅推理效率低,还易引入语义合理但因果虚假的奖励成分,导致优化无效。为此,我们提出因果奖励世界模型(CRWM),通过在多任务交互数据上离线预训练,显式建模候选奖励组件与目标任务物理变量之间的因果拓扑关系。采用粗到细的预训练策略,引入联合优化模块,结合显式机制解耦与置信度感知软融合,利用微观轨迹细化粗粒度结构先验,构建稳健且可解释的因果骨架。推理时,LLM将CRWM作为与任务无关的因果先验,约束奖励生成,实现零样本奖励函数设计。大量实验表明,CRWM无需反馈驱动的奖励优化即可生成可执行奖励函数,在复杂连续控制基准测试中显著降低新技能获取的设计延迟,性能达到或超越现有最佳水平,并展现出对未见任务和多样化机器人形态的强大泛化能力。

原文摘要 · Abstract (English)

Automated Reward Design (ARD) aims to replace manual reward engineering in reinforcement learning with language-driven reward function synthesis. However, existing approaches based on large language models (LLMs) remain inherently correlation-driven, relying on iterative environmental feedback to refine reward hypotheses for each specific task. This paradigm not only results in inefficient reasoning but also makes LLMs susceptible to semantically plausible yet causally spurious reward components, leading to ineffective optimization. To address these limitations, we propose the Causal Reward World Model (CRWM), which explicitly models the causal topological relationships between candidate reward components and task-targeted physical variables through offline pre-training on multi-task interaction data. Based on a coarse-to-fine pre-training strategy, we introduce a joint optimization module that integrates Explicit Mechanism Decoupling with Confidence-Aware Soft Fusion to refine coarse structural priors using micro-level trajectories, thereby constructing a robust and interpretable causal skeleton. During inference, LLMs leverage CRWM as a task-irrelevant causal prior to constrain the reward generation, enabling zero-shot reward function design. Our work opens up a new white-box paradigm for the ARD problem. Extensive experiments on complex continuous control benchmarks demonstrate that CRWM generates executable reward functions without feedback-driven reward refinement, significantly reducing the design latency for acquiring new robotic skills while matching or surpassing state-of-the-art performance, and further exhibits strong generalization capabilities across unseen tasks and diverse robotic embodiments.

强化学习因果建模零样本机器人技能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。