用权重空间建模视频,实现高效可编辑的虚拟世界生成
Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement

- 将系统状态编码为神经网络权重,直接解析渲染,无需解码器
- 在单张消费级显卡上运行,参数量仅4000万,支持零样本超分辨率
- 自动分离背景、前景和运动,可单独编辑内容或动态而不互相干扰
在海量未标注视频上训练世界模型是实现完全自主智能的关键步骤。然而,现有方法将原始像素编码为不透明的隐空间,并依赖复杂的解码器进行重建,导致计算成本高且难以解释。我们提出NOVA框架,将系统状态表示为辅助坐标基隐式神经网络(INR)的权重与偏置。该结构化表示可解析渲染,消除解码瓶颈,同时具备紧凑性、可移植性和零样本超分辨率能力。此外,与多数隐动作模型类似,NOVA可通过动作匹配目标蒸馏为上下文相关的视频生成器。令人惊讶的是,无需额外损失或对抗性目标,NOVA能自动分离场景中的背景、前景及帧间运动成分,实现内容与动态的独立编辑而不相互影响。我们在多个挑战性数据集上验证了该框架,在单个消费级GPU上以约4000万参数运行,实现了强可控预测。结构化表示如INR不仅深化了对隐动态的理解,也为沉浸式、可定制的虚拟体验铺平道路。
原文摘要 · Abstract (English)
Training world models on vast quantities of unlabelled videos is a critical step toward fully autonomous intelligence. However, the prevailing paradigm of encoding raw pixels into opaque latent spaces and relying on heavy decoders for reconstruction leaves these models computationally expensive and uninterpretable. We address this problem by introducing NOVA, a world modelling framework that represents the system state as the weights and biases of an auxiliary coordinate-based implicit neural representation (INR). This structured representation is analytically rendered, which eliminates the decoder bottleneck while conferring compactness, portability, and zero-shot super-resolution. Furthermore, like most latent action models, NOVA can be distilled into a context-dependent video generator via an action-matching objective. Surprisingly, without resorting to auxiliary losses or adversarial objectives, NOVA can disentangle structural scene components such as background, foreground, and inter-frame motion, enabling users to edit either content or dynamics without compromising the other. We validate our framework on several challenging datasets, achieving strong controllable forecasting while operating on a single consumer GPU at $\sim$40M parameters. Ultimately, structured representations like INRs not only enhance our understanding of latent dynamics but also pave the way for immersive and customisable virtual experiences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。