对比重建与语义潜空间,发现语义空间更适合作为机器人世界模型基础。
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models

- 用六种潜空间在固定协议下训练动作条件扩散模型
- 语义编码器在规划与策略性能上全面优于重建编码器
- 强调视觉保真度不足以为世界模型选型依据,适合机器人控制研究者
基于世界模型的策略评估是通过在动作条件视频扩散模型中回放候选动作来模拟真实机器人控制的有效方法。随着模型越来越多采用潜扩散建模(LDM),选择合适的潜空间变得至关重要。当前主流使用以像素重建为目标的自编码潜空间(如VAE),但近期研究指出,具备语义对齐能力的预训练编码器更具优势。我们系统评估了六种重建与语义编码器在固定协议下的表现,使用BridgeV2数据集训练世界模型变体,并验证了在高维表示空间中,无论是否压缩维度,均可实现有效训练。提出三个评估轴:视觉保真度、规划与下游策略性能、潜表示质量。结果表明,仅依赖视觉保真度不足以选型世界模型。尽管重建编码器(如VAE和Cosmos)在像素级指标上表现优异,但语义编码器(如V-JEPA 2.1、Web-DINO、SigLIP 2)在所有模型规模下均在其他两个维度上表现更优,证明语义潜空间是更优的政策相关机器人扩散世界模型基础。
原文摘要 · Abstract (English)
World model-based policy evaluation is a practical proxy for testing real-world robot control by rolling out candidate actions in action-conditioned video diffusion models. As these models increasingly adopt latent diffusion modeling (LDM), choosing the right latent space becomes critical. While the status quo uses autoencoding latent spaces like VAEs that are primarily trained for pixel reconstruction, recent work suggests benefits from pretrained encoders with representation-aligned semantic latent spaces. We systematically evaluate these latent spaces for action-conditioned LDM by comparing six reconstruction and semantic encoders to train world model variants under a fixed protocol on BridgeV2 dataset, and show effective world model training in high-dimensional representation spaces with and without dimension compression. We then propose three axes to assess robotic world model performance: visual fidelity, planning and downstream policy performance, and latent representation quality. Our results show visual fidelity alone is insufficient for world model selection. While reconstruction encoders like VAE and Cosmos achieve strong pixel-level scores, semantic encoders such as V-JEPA 2.1 (strongest overall on policy), Web-DINO, and SigLIP 2 generally excel across the other two axes at all model scales. Our study advocates semantic latent space as stronger foundation for policy-relevant robotics diffusion world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。