用3D视频生成构建可交互的4D世界,支持实时多模态操作
3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation
- 通过四模块集成与视觉聚焦渲染,将图文转为连贯4D场景
- 采用眼动聚焦策略实现高效实时交互,支持用户自适应探索
- 适合虚拟现实、数字孪生等需动态可视化场景的研究者
我们提出3D4D,一个结合WebGL与Supersplat渲染的交互式4D可视化框架。该框架通过四个核心模块,将静态图像与文本转化为连贯的4D场景,并采用眼动聚焦渲染策略,实现高效、实时的多模态交互,支持用户自适应地探索复杂4D环境。项目主页与代码已公开于https://yunhonghe1021.github.io/NOVA/。
原文摘要 · Abstract (English)
We introduce 3D4D, an interactive 4D visualization framework that integrates WebGL with Supersplat rendering. It transforms static images and text into coherent 4D scenes through four core modules and employs a foveated rendering strategy for efficient, real-time multi-modal interaction. This framework enables adaptive, user-driven exploration of complex 4D environments. The project page and code are available at https://yunhonghe1021.github.io/NOVA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。