arXiv:2509.21073cs.RO2025-09被引 2

用归一化流构建双手操作策略,推理快且能量化不确定性。

Normalizing Flows are Capable Models for Bi-manual Visuomotor Policy

  • 基于条件归一化流建模动作序列分布,支持单次生成采样。
  • 在仿真和真实机器人上成功率超基线,训练更快、推理延迟更低。
  • 适合需要实时响应与不确定性感知的机器人控制场景。

通用机器人领域近年来采用强大的概率扩散模型来学习复杂的具身行为,但现有模型常伴随显著的计算开销和无法量化输出不确定性的根本缺陷。我们提出基于归一化流的双臂视觉-运动策略(NF-P),学习动作序列的条件密度,并支持单次生成采样与可计算的似然值。利用该特性,我们设计两种推理时优化策略:随机批量选择(从采样候选中选最高似然轨迹)和梯度精炼(直接沿对数似然梯度上升以提升动作质量)。在仿真与真实机器人实验中,NF-P表现优于基线,同时具备更快训练速度和更低推理延迟。结果表明,归一化流是具竞争力且计算高效的视觉-运动策略,尤其适用于实时、不确定性感知的机器人控制。

原文摘要 · Abstract (English)

The field of general-purpose robotics has recently embraced powerful probabilistic diffusion-based models to learn the complex embodiment behaviours. However, existing models often come with significant trade-offs, namely high computational costs for inference and a fundamental inability to quantify output uncertainty. We introduce Normalizing Flows Policy (NF-P), a conditional normalizing flow-based visuomotor policy for bi-manual manipulation. NF-P learns a conditional density over action sequences and enables single-pass generative sampling with tractable likelihood computation. Using this property, we propose two inference-time optimization strategies: Stochastic Batch Selection, which selects the highest-likelihood trajectory among sampled candidates, and Gradient Refinement, which directly ascends the log-likelihood to improve action quality. In both simulation and real robot experiments, NF-P achieves promising success rates compared to the baseline. In addition to improved task performance, NF-P demonstrates faster training and lower inference latency. These results establish normalizing flows as a competitive and computationally efficient visuomotor policy, particularly for real-time, uncertainty-aware robotic control.

视觉运动归一化流双臂操作不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。