arXiv:2604.24764cs.CV2026-04被引 8

用强化学习提升文本生成视频的3D一致性,不改架构也能保持画质。

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation

论文配图:World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
图 1 · 摘自论文原文
  • 通过强化学习让视频生成符合3D结构约束,无需修改模型架构。
  • 在多个数据集上3D一致性指标提升显著,同时保持原始视觉质量。
  • 适合关注视频生成真实感与世界模拟的研究者或开发者。

近期视频基础模型虽具备出色的视觉合成能力,但常出现几何不一致问题。现有方法通过架构修改注入3D先验,但计算开销大且难以扩展。本文提出World-R1框架,通过强化学习将视频生成对齐3D约束。为此,我们构建了一个专用于世界模拟的纯文本数据集。利用Flow-GRPO算法,基于预训练3D基础模型和视觉语言模型的反馈优化模型,以增强结构连贯性,且不改变底层架构。进一步采用周期性解耦训练策略,在刚性几何一致性与动态场景流畅性间取得平衡。大量实验表明,该方法显著提升3D一致性,同时保留基础模型的原始视觉质量,有效弥合视频生成与可扩展世界模拟之间的差距。

原文摘要 · Abstract (English)

Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high computational costs and limit scalability. We propose World-R1, a framework that aligns video generation with 3D constraints through reinforcement learning. To facilitate this alignment, we introduce a specialized pure text dataset tailored for world simulation. Utilizing Flow-GRPO, we optimize the model using feedback from pre-trained 3D foundation models and vision-language models to enforce structural coherence without altering the underlying architecture. We further employ a periodic decoupled training strategy to balance rigid geometric consistency with dynamic scene fluidity. Extensive evaluations reveal that our approach significantly enhances 3D consistency while preserving the original visual quality of the foundation model, effectively bridging the gap between video generation and scalable world simulation.

视频生成3D一致性强化学习世界模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。