提出可分离结构与内容的潜空间动作模型,兼顾抽象动作与生成质量。
DiLA: Disentangled Latent Action World Models

- 通过解耦结构与内容路径,实现潜空间动作的语义化与连续表示
- 在视频生成、动作迁移等任务上优于现有方法,生成质量高且可解释性强
- 适合自监督世界模型、视频生成与视觉规划方向的研究者
潜空间动作模型(LAMs)能从无标签视频中学习世界模型,通过推断连续帧间的抽象动作实现。然而,这类方法面临动作抽象与生成保真度之间的根本权衡。现有方法通常采用两阶段训练或仅预测光流来规避此问题。本文提出DiLA,一种基于内容-结构解耦的新型潜空间动作世界模型,旨在通过解耦机制解决该权衡。核心洞察在于:潜空间动作学习中的预测瓶颈自然驱动解耦,迫使模型将空间布局编码至结构路径,同时将视觉细节交由独立的内容路径生成。这种协同作用构建了连续、语义清晰的潜动作空间,不牺牲生成质量。DiLA在视频生成质量、动作迁移、视觉规划和流形可解释性方面均取得更优结果,成为统一实现高层次动作抽象与高保真生成的框架,推动自监督世界模型学习的前沿发展。
原文摘要 · Abstract (English)
Latent Action Models (LAMs) enable the learning of world models from unlabeled video by inferring abstract actions between consecutive frames. However, LAMs face a fundamental trade-off between action abstraction and generation fidelity. Existing methods typically circumvent this issue by using two-stage training with pre-trained world models or by limiting predictions to optical flow. In this paper, we introduce DiLA, a novel Disentangled Latent Action world model that aims to resolve this trade-off via content-structure disentanglement. Our key insight is that disentanglement and latent action learning are co-evolving: the predictive bottleneck inherent in latent action learning serves as a driving force for disentanglement, compelling the model to distill spatial layouts into the structure pathway while offloading visual details to a separate content pathway for generation. This synergy yields a continuous, semantically structured latent action space without compromising generative quality. DiLA achieves superior results in video generation quality, action transfer, visual planning, and manifold interpretability. These findings establish DiLA as a unified framework that simultaneously achieves high-level action abstraction and high-fidelity generation, advancing the frontier of self-supervised world model learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。