为机器人操作设计的新型视频自编码器,让动作更可控、压缩更高效
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation

- 双编码器单解码器结构,自动分离机械臂动作与背景
- 压缩率高且动作控制精度提升2dB PSNR
- 适合需要精细动作控制的机器人学习任务
基于扩散模型的具身学习近年取得显著进展,但现有模型依赖适用于自然场景的变分自编码器(VAE),未能考虑具身操作的独特性,导致潜在表示冗余且难以控制,影响训练效率和机器人精准操控。为此,我们提出EmbodiedVAE,一种专为机器人操作世界模型设计的视频自编码器。该模型采用双编码器-单解码器架构,结合非对称时空压缩模块,可自动将机械臂运动与环境背景解耦,实现紧凑表示并提供显式具身潜在变量以支持细粒度动作控制。为进一步保持动作时序一致性,引入基于最优传输的一致性模块,显式强化运动保真度与帧间连贯性。大量实验表明,所提EmbodiedVAE在保持高压缩率的同时,重建质量更优,相较当前最优视频VAE,在机器人操作场景中平均提升2dB PSNR,实现更精确的动作控制。
原文摘要 · Abstract (English)
Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable performance, existing LDMs predominantly rely on Variational Autoencoders (VAEs) optimized for natural scenes while failing to account for the unique characteristics of embodied manipulation scenarios, yielding latent representations that are neither compact nor controllable, thereby hindering efficient training of LDMs and precise robotic control. To solve this problem, we present EmbodiedVAE, a novel video VAE that provides compact yet controllable latent representations tailored for the robotic manipulation world models. Specifically, EmbodiedVAE adopts a dual-encoder, single-decoder architecture with an asymmetric spatio-temporal compression module, which automatically disentangles the robot arm's motion from background environment, resulting in overall compactness while providing explicit embodied latent to support fine-grained action control. To further preserve the temporal consistency of learned robotic motion latent, we introduce an optimal-transport-based consistency module that explicitly enforces motion fidelity and inter-frame coherence. Extensive experiments demonstrate that our proposed EmbodiedVAE achieves superior reconstruction quality with high compression rate, while enabling more precise action control in robotic manipulation scenarios with an average of 2dB PSNR improvement over state-of-the-art video VAEs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。