实时生成无限长3D一致视频,保持场景稳定与物理合理性。
Endless World: Real-Time 3D-Aware Long Video Generation
- 用条件自回归训练确保新内容与已有帧对齐,支持无限续播。
- 单卡实时推理,生成视频在视觉与空间一致性上表现优异。
- 适合需要持续生成动态3D场景的应用,如虚拟世界、元宇宙。
生成具有稳定3D结构的长时序连贯视频仍是重大挑战,尤其在流式场景中。为此,我们提出Endless World,一个支持无限、3D一致视频实时生成的框架。为实现无限生成,引入条件自回归训练策略,使新生成内容与已有视频帧对齐,既保留长程依赖又保持计算高效,可在单张GPU上实时推理且无需额外训练开销。此外,集成全局3D感知注意力,提供跨时间的连续几何引导;3D注入机制强制保证整个长序列中的物理合理性和几何一致性,有效解决长时域动态场景合成的关键难题。大量实验表明,Endless World能生成长时、稳定且视觉连贯的视频,在视觉保真度与空间一致性上达到或优于现有方法。项目已开源:https://bwgzk-keke.github.io/EndlessWorld/
原文摘要 · Abstract (English)
Producing long, coherent video sequences with stable 3D structure remains a major challenge, particularly in streaming scenarios. Motivated by this, we introduce Endless World, a real-time framework for infinite, 3D-consistent video generation.To support infinite video generation, we introduce a conditional autoregressive training strategy that aligns newly generated content with existing video frames. This design preserves long-range dependencies while remaining computationally efficient, enabling real-time inference on a single GPU without additional training overhead.Moreover, our Endless World integrates global 3D-aware attention to provide continuous geometric guidance across time. Our 3D injection mechanism enforces physical plausibility and geometric consistency throughout extended sequences, addressing key challenges in long-horizon and dynamic scene synthesis.Extensive experiments demonstrate that Endless World produces long, stable, and visually coherent videos, achieving competitive or superior performance to existing methods in both visual fidelity and spatial consistency. Our project has been available on https://bwgzk-keke.github.io/EndlessWorld/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。