开源143万帧视频数据集,支持多源数据统一训练长时序世界模型。
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

- 构建可重构多源数据引擎,统一10个数据集的视觉与相机信息
- 用5秒短序列训练出能实时交互数分钟至数小时的因果模型
- 支持不同视频模型架构无缝接入,提升可复现性与扩展性
我们提出SolarWM,一个从数据准备到长时序推理全程开源的基础框架,用于构建交互式视频世界模型。跨异构数据源与视频骨干网络的训练面临挑战:数据集在时间尺度、相机几何、视觉质量、运动特征和描述风格上差异显著,而视频生成器采用不同表示与架构。直接混合数据与特定模型实现会导致监督不一致,难以复现与比较。SolarWM通过可重构的多源数据引擎与骨干原生适配框架解决此耦合问题。该引擎将来自10个数据集的143万条标准视频片段转化为统一帧对齐的数据契约,涵盖视觉观测、度量相机几何、字幕、质量元数据、筛选决策与来源信息,并将源处理与混合构建解耦。在共享相机条件、训练与推理接口下,我们基于Wan2.2、LTX-2.5和MiniMax-H3实现了四个5B至33B参数量的模型,同时保留其原生表示与目标。采用统一三阶段方案:双向适配、教师强制自回归初始化与分布匹配蒸馏。最终模型仅用5秒序列训练即可实现在分钟至小时级滚动中的实时交互。通过发布数据、流程、配方、权重与框架,SolarWM为交互式世界模型研究提供了可复现且可扩展的基础。
原文摘要 · Abstract (English)
We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ in temporal scale, camera geometry, visual quality, motion, and captioning styles, while video generators use distinct representations and architectures. Naive data mixing and model-specific implementations therefore produce inconsistent supervision and make results difficult to reproduce and compare. SolarWM addresses this coupling with a reconfigurable multi-source data engine and a backbone-native adaptation framework. The engine converts 1.43 million canonical clips from 10 datasets into a unified, frame-aligned contract covering visual observations, metric camera geometry, captions, quality metadata, selection decisions, and provenance, while decoupling source processing from mixture construction. Under shared camera-conditioning, training, and inference interfaces, we instantiate four 5B--33B models based on Wan2.2, LTX-2.5, and MiniMax-H3 while preserving their native representations and objectives. A unified three-stage recipe combines bidirectional adaptation, teacher-forced autoregressive initialization, and distribution matching distillation. The resulting causal models enable real-time interaction over rollouts ranging from minutes to hours after being trained on only 5s sequences. By releasing the resulting data, pipeline, recipes, weights, and framework, SolarWM provides a reproducible and extensible foundation for interactive world-model research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。