arXiv:2605.23993cs.CVcs.AI2026-05被引 1

轻量级视频预测框架,便于研究世界模型设计选择

Nano World Models: A Minimalist Implementation of Future Video Prediction

论文配图:Nano World Models: A Minimalist Implementation of Future Video Prediction
图 1 · 摘自论文原文
  • 基于扩散强制的极简代码实现,统一多种生成目标与配置
  • 在控制环境、游戏和真实机器人数据上验证不同组件影响
  • 适合希望复现或对比世界模型组件的研究者使用

世界模型已成为支持生成、规划与决策的预测模拟器的核心范式。尽管工业级交互式视频生成进展迅速,但学术界仍缺乏紧凑、可复现且易扩展的世界模型实现。我们提出 Nano World Models,一个以扩散强制为核心的未来视频预测极简代码库。该框架统一了生成目标、模型规模、动作条件机制、潜在观测空间、数据集、评估协议及长时程回放流程,使原本分散实现中纠缠的组件得以可控研究。通过在简单控制环境、游戏仿真和真实机器人数据上的实验,我们分析了预测参数化、架构规模、动作注入、采样预算与领域复杂度对视频预测质量及自回归回放行为的影响。通过发布代码、配置文件、评估脚本和预训练检查点,Nano World Models旨在为开放、可复现、科学的世界模型研究提供紧凑而可扩展的实验基础。

原文摘要 · Abstract (English)

World models have become a central paradigm for learning predictive simulators that support generation, planning, and decision-making. Yet, despite rapid progress in industry-scale interactive video generation, the broader research community still lacks compact, reproducible, and easily extensible implementations for studying the design choices underlying modern world models. We introduce Nano World Models, a minimalist codebase for future video prediction centered around diffusion forcing. Nano World Models provides a unified interface for generative objectives, model scales, action-conditioning mechanisms, latent observation spaces, datasets, evaluation protocols, and long-horizon rollout procedures. This design enables controlled studies of world-modeling components that are often entangled across separate implementations. Through experiments across simple control environments, game simulation, and real-robot data, we examine how prediction parameterization, architecture scale, action injection, sampling budget, and domain complexity affect video prediction quality and autoregressive rollout behavior. By releasing code, configurations, evaluation scripts, and pretrained checkpoints, Nano World Models aims to provide a compact yet extensible experimental substrate for open, reproducible, and scientific world-model research.

视频预测世界模型扩散模型可复现研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。