arXiv:2605.22123cs.RO2026-05

从少量示范中学习不变奖励,让机器人在真实世界跨场景通用

Beyond Pixels: Learning Invariant Rewards for Real-World Robotics From a Few Demonstrations

论文配图:Beyond Pixels: Learning Invariant Rewards for Real-World Robotics From a Few Demonstrations
图 1 · 摘自论文原文
  • 通过发现任务级行为不变性替代像素拟合
  • 仅需5次示范即可生成可泛化的奖励函数
  • 零样本适配新位置/视角/物体,适合真实机器人部署

在开放世界操作任务中,同一任务可能因物体实例、位置和相机视角不同而呈现多种变体。现有基于视觉的奖励模型常记忆特定像素分布,难以泛化。本文提出框架,仅需5次示范即可学习不变的符号化奖励函数。核心思想是从视觉差异中发现任务级不变特性:在不同视觉条件下保持恒定的行为规律。该框架包含两个耦合组件:一是编码任务策略与物理约束的结构化奖励形式,保证最优策略不变性;二是混合符号-数值方法,从示范中提炼这些不变性,无需在线交互。在八项Meta-World任务和三项Franka机械臂任务上验证,本方法显著提升过程对齐与策略回放排序能力,加速下游策略学习。三次真实世界分布外实验表明,同一奖励函数可零样本泛化至位置、视角和物体变化,实现单一奖励表示在多样任务变体中的复用。

原文摘要 · Abstract (English)

Designing reward functions that generalize beyond controlled laboratory settings remains a fundamental challenge in reinforcement learning for robotics. In open-world manipulation problems, a single task can appear in numerous variants through different object instances, positions, and camera viewpoints. Recent vision-based reward models tend to memorize specific pixel distributions and fail to generalize beyond their training conditions. To address this, we propose a framework that learns invariant symbolic reward functions from as few as five demonstrations. The insight is to shift from visual feature-fitting to the discovery of behavioral invariants: task-level properties that remain constant across diverse visual instantiations. The framework has two coupled components: a structural reward formulation that encodes task-level strategies and physical constraints while preserving optimal policy invariance, and a hybrid symbolic-numerical procedure that distills these invariants from demonstrations without online interaction. Experiments on eight Meta-World tasks and three Franka manipulation tasks demonstrate that our method achieves stronger process alignment and policy rollout ranking abilities compared to baselines, accelerating downstream policy learning. Three real-world out-of-distribution experiments further show that the same learned reward generalizes zero-shot to position, viewpoint, and object variations, enabling a single reward representation to be reused across diverse task variants in practice.

机器人强化学习奖励学习泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。