WorldDiT用统一模型同时生成动作与建模视觉世界,无需大VLM即可实现强性能。
WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

- 统一扩散变压器架构,同步生成动作与预测未来视觉补丁
- 在4个LIBERO仿真任务中达到参数量与成功率的帕累托最优
- 适合追求轻量级高效率机器人控制的研究者参考
许多近期机器人策略通过使用大型预训练视觉语言模型(VLM)作为动作主干来提升控制能力。我们提出WorldDiT,一种统一的扩散变换器架构,将动作生成与视觉世界建模相结合,在无需大型预训练VLM动作主干的情况下实现了强大性能。训练时,单一扩散变换器同时生成连续动作片段,并从未来的相机帧中预测归一化RGB补丁。在四个LIBERO仿真套件中,WorldDiT位于报告的帕累托前沿,其总模型参数与平均成功率均优于现有方法。这些结果为未来扩展研究提供了强大的亚十亿参数基线。
原文摘要 · Abstract (English)
Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone. During training, a single diffusion transformer generates continuous action chunks and predicts normalized RGB patch targets from future camera frames. Across four LIBERO simulation suites, WorldDiT lies on the reported Pareto frontier for total model parameters and mean success among methods reporting all four suites. These results provide a strong sub-billion-parameter baseline for future scaling studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。