提出连续动作表示框架,让机器人任务更稳定高效。
Trajectory-Level Continuous Action Representation for Robotic Manipulation

- 用连续潜变量表示整个操作轨迹,不依赖固定时间参数
- 在多个数据集上提升成功率,最高达18%以上
- 适合不同控制频率的长程机器人任务
我们提出CAT,一种用于机器人操作的轨迹级连续动作表示框架。现有视觉-运动系统常将动作表示与控制频率绑定或依赖预设的时间参数化,导致高采样率下表示冗余,并限制关键运动建模。CAT将固定时长的动作轨迹编码为一组连续潜变量,并引入频率感知的位置编码,建立共享的时间坐标系以保证跨频率的时序一致性。轨迹级正则化进一步稳定潜空间表示。该方法避免了随时间步密度增长的表示膨胀,且无需预设时间参数。在LIBERO、MimicGen及真实世界长程操作任务上的系统级评估表明,基于CAT的策略在匹配训练设置下持续优于多种基于向量量化和连续的基线方法。无论模型主干或控制频率如何,CAT均显著提升成功率。结果凸显了轨迹级连续动作建模在可扩展机器人操作中的优势。
原文摘要 · Abstract (English)
We propose CAT, a trajectory-level continuous action representation framework for robotic manipulation. Existing visuomotor systems often entangle action representation with control frequency or rely on fixed temporal parameterizations. This leads to representational redundancy at high sampling rates and limits the modeling of critical motion. CAT instead encodes action trajectories within a fixed real-time interval into a set of continuous latent tokens. To ensure temporal consistency across varying control frequencies, we further incorporate a frequency-aware positional encoding that establishs a shared temporal coordinate system. Trajectory-level regularization further stabilizes the latent representation. This approach prevents representation growth with timestep density and avoids reliance on predefined temporal parameterizations. Extensive system-level evaluations on LIBERO, MimicGen, and real-world long-horizon manipulation tasks demonstrate that CAT-based policies consistently outperform both competitive VQ-based and continuous visuomotor baselines under matched training settings. Across various model backbones and control frequencies, CAT consistently improves success rates. These results highlight the advantages of trajectory-level continuous action modeling for scalable robotic manipulation across varying control rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。