arXiv:2606.17408cs.ROcs.CV2026-06

用可学习的初始分布提升机器人生成策略的性能。

Where Should Action Generation Begin? A Learnable Source Prior for Generative Robot Policies

论文配图:Where Should Action Generation Begin? A Learnable Source Prior for Generative Robot Policies
图 1 · 摘自论文原文
  • 用本体感知条件的对角高斯分布替代标准高斯作为动作生成起点。
  • 在15个任务上平均成功率81.6%,比基线高出6.5至25.5个百分点。
  • 轻量设计,加速收敛,适合真实场景部署,通用性强。

生成式机器人策略通常从与观测无关的标准高斯分布开始动作生成,而源分布的选择尚未被充分探索。本文提出可学习源先验(LeaP),以本体感知条件的对角高斯分布取代标准高斯分布,作为动作片段的初始分布。该分布由一个轻量级MLP参数化,联合预测均值与状态自适应方差,同时保持下游生成器架构和推理求解器不变。此设计提供有观测依据且带随机性的初始化,使生成器更专注于精细动作修正而非从无信息噪声源中迁移样本。在15个RoboTwin操控任务上,LeaP平均成功率达81.6%,优于四种代表性基线——包括确定性源方法、无先验对照组及扩散桥策略——提升6.5至25.5个百分点。相同先验同时改善流匹配与扩散桥生成器,参数更少、收敛更快。优势延续至真实世界部署,表现最优。结果表明,源分布是生成式机器人策略中独立且可复用的设计维度,与生成动力学选择互补。

原文摘要 · Abstract (English)

Generative robot policies typically begin action generation from an observation-independent standard Gaussian distribution, leaving the choice of source distribution underexplored. This work asks a simple question: where should action generation begin? We propose LeaP, a Learnable source Prior that replaces the standard Gaussian with a proprioception-conditioned diagonal Gaussian over action chunks. Parameterized by a lightweight MLP, LeaP jointly predicts the mean and state-adaptive variance of the source distribution, while keeping the downstream generator architecture and inference solver unchanged. This design provides an observation-informed yet stochastic initialization, allowing the generator to focus on precise action refinement rather than transporting samples from an uninformed noise source. On 15 RoboTwin manipulation tasks, LeaP achieves an average success rate of 81.6%, outperforming four representative baselines -- including deterministic-source methods, a no-prior counterpart, and a diffusion-bridge policy -- by 6.5 to 25.5 percentage points. The same prior consistently improves both flow-matching and diffusion-bridge generators, while using fewer parameters and converging faster. The advantage carries over to real-world deployment, where LeaP attains the best performance. These results suggest that the source distribution is an independent and reusable design axis for generative robot policies, complementary to the choice of generative dynamics.

机器人生成策略先验设计强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。