arXiv:2602.16229cs.LG2026-02被引 5

将复杂视频中的多个主体动作解耦建模,提升生成与控制精度。

Factored Latent Action World Models

  • 将场景分解为独立因子,每个因子独立学习潜变量动作
  • 在多主体数据集上预测准确率和生成质量显著优于单一潜变量模型
  • 适合需要精细控制的视频生成与策略学习任务

从无动作视频中学习潜变量动作已成为扩展可控世界模型学习的有效范式。潜变量动作为用户迭代生成和操控视频提供了自然接口。然而,现有方法大多依赖于整体式的逆向与前向动力学模型,仅学习单一潜变量动作来控制整个场景,在多个实体同时活动的复杂环境中表现不佳。本文提出因子化潜变量动作模型(FLAM),采用因子化动力学框架,将场景分解为独立因子,每个因子独立推断其潜变量动作并预测下一时刻的因子值。该结构能更精确建模复杂多主体动态,在无动作视频设置下相比整体模型提升了视频生成质量。基于仿真和真实多主体数据集的实验表明,FLAM在预测精度和表征质量上均优于先前方法,并有助于下游策略学习,验证了因子化潜变量动作模型的优势。

原文摘要 · Abstract (English)

Learning latent actions from action-free video has emerged as a powerful paradigm for scaling up controllable world model learning. Latent actions provide a natural interface for users to iteratively generate and manipulate videos. However, most existing approaches rely on monolithic inverse and forward dynamics models that learn a single latent action to control the entire scene, and therefore struggle in complex environments where multiple entities act simultaneously. This paper introduces Factored Latent Action Model (FLAM), a factored dynamics framework that decomposes the scene into independent factors, each inferring its own latent action and predicting its own next-step factor value. This factorized structure enables more accurate modeling of complex multi-entity dynamics and improves video generation quality in action-free video settings compared to monolithic models. Based on experiments on both simulation and real-world multi-entity datasets, we find that FLAM outperforms prior work in prediction accuracy and representation quality, and facilitates downstream policy learning, demonstrating the benefits of factorized latent action models.

世界模型潜变量动作多主体建模视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。