arXiv:2606.21139cs.ROcs.AI2026-06

将动作的幅度与模式分离,提升机器人策略学习效果

PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning

论文配图:PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning
图 1 · 摘自论文原文
  • 用径向结构分离动作幅度(半径)与模式(方向)
  • 在超球空间中建模,大跨度动作支持更多模式变化
  • 在仿真和真实机器人上均优于现有方法

隐式动作预训练从观察对中学习视觉变化表征,但现有方法通常将每个转换编码为单一无结构表示,混淆了动作幅度与动作模式。我们提出具有径向结构的极坐标隐式动作(PoLAR),强制隐式动作具有径向-方向结构:半径编码动作幅度,方向保留动作模式。PoLAR 使用两个观测之间的时序偏移作为动作幅度的弱代理,使时间间隔较大的观测对对应更大的半径。我们在双曲空间中实现该结构,其随半径增大而体积扩张的特性天然适配大动作幅度下的多样化模式。在任务内及大规模预训练设置下,PoLAR 在仿真和真实机器人实验中均显著提升下游策略性能,优于现有的隐式动作基线和强基准视觉语言动作模型(VLAs)。结果表明,隐式动作空间的几何设计是将视觉预训练迁移到下游机器人策略学习的关键因素。

原文摘要 · Abstract (English)

Latent action pretraining learns representations of visual change from pairs of observations, but existing methods typically encode each transition as a single unstructured representation that entangles transition extent and transition mode. We introduce Polar Latent Actions with Radial structure (PoLAR), which imposes a radial-direction structure on latent actions, encouraging radius to encode transition extent and direction to retain transition mode. PoLAR uses temporal offset between two observations as a weak proxy for transition extent, encouraging latent action from observation pairs separated by larger temporal gaps to occupy larger radii. We instantiate this structure in hyperbolic space, whose expanding volume with radius offers a natural fit for more diverse transition modes at larger extents. Across in-task and large-scale pretraining settings, PoLAR improves downstream policy performance in simulation and real-world robot experiments, outperforming latent action baselines and strong pretrained VLAs. These results suggest that the geometry of the latent action space is an important design choice for transferring visual pretraining to downstream robot policy learning.

机器人策略隐式动作几何建模迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。