arXiv:2501.14174cs.CVcs.AI2025-01ICLR被引 5

从视频中自动分解物体与属性,实现无辅助数据的未来画面重组合生成。

Dreamweaver: Learning Compositional World Models from Pixels

  • 用递归块槽单元分解视频为对象与属性,构建层次化表示。
  • 在多个数据集上超越当前最优基线,动态与静态概念解耦更优。
  • 支持属性重组生成新视频,适合需要创造性生成的场景。

人类天生能将世界感知分解为物体及其属性(如颜色、形状、运动模式),并借此重构想象未来。然而,人工智能系统在不依赖文本、掩码或边界框等辅助信息的情况下,仍难以建模视频的组成结构并生成未见的重组未来。本文提出Dreamweaver,一种神经架构,可从原始视频中发现分层且组合式的表征,并生成组合式未来模拟。其核心是新型的递归块槽单元(RBSU),用于将视频分解为构成对象与属性;同时采用多未来帧预测目标,更有效捕捉动态与静态概念的解耦表示。实验表明,在DCI框架下,该模型在多个数据集上优于现有最先进基线。此外,模型模块化的概念表示支持组合式想象,可通过重组已见对象的属性生成全新视频。

原文摘要 · Abstract (English)

Humans have an innate ability to decompose their perceptions of the world into objects and their attributes, such as colors, shapes, and movement patterns. This cognitive process enables us to imagine novel futures by recombining familiar concepts. However, replicating this ability in artificial intelligence systems has proven challenging, particularly when it comes to modeling videos into compositional concepts and generating unseen, recomposed futures without relying on auxiliary data, such as text, masks, or bounding boxes. In this paper, we propose Dreamweaver, a neural architecture designed to discover hierarchical and compositional representations from raw videos and generate compositional future simulations. Our approach leverages a novel Recurrent Block-Slot Unit (RBSU) to decompose videos into their constituent objects and attributes. In addition, Dreamweaver uses a multi-future-frame prediction objective to capture disentangled representations for dynamic concepts more effectively as well as static concepts. In experiments, we demonstrate our model outperforms current state-of-the-art baselines for world modeling when evaluated under the DCI framework across multiple datasets. Furthermore, we show how the modularized concept representations of our model enable compositional imagination, allowing the generation of novel videos by recombining attributes from previously seen objects. cun-bjy.github.io/dreamweaver-website

视频生成组合建模自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。