构建400万帧动态数据集,提升真实场景下生成渲染的精度与一致性。
Generative World Renderer
- 用双屏拼接采集400万帧高保真视频数据,含五通道G-buffer。
- 逆向渲染模型在新数据上泛化能力显著提升,生成视频更逼真。
- 适合游戏/影视生成、跨域渲染研究者使用。
将生成式逆向与正向渲染扩展至真实世界场景,受限于现有合成数据集的现实感和时间连贯性不足。为弥合这一持续存在的领域差距,我们引入一个大规模动态数据集,源自视觉复杂的3A游戏。通过一种新颖的双屏拼接采集方法,我们提取了400万帧连续帧(720p/30 FPS)的同步RGB及五通道G-buffer数据,覆盖多样场景、视觉效果与环境,包括恶劣天气和运动模糊变体。该数据集独特地推进双向渲染:支持真实环境中几何与材质的鲁棒分解,并促进基于G-buffer的高保真视频生成。此外,为在无真值条件下评估逆向渲染的真实世界性能,我们提出一种基于视觉语言模型的评估协议,衡量语义、空间与时间一致性。实验表明,基于本数据微调的逆向渲染器在跨数据集泛化和可控生成方面表现优异,且该评估方法与人类判断高度相关。结合我们的工具包,正向渲染器支持用户通过文本提示编辑3A游戏的风格,仅需G-buffers输入。
原文摘要 · Abstract (English)
Scaling generative inverse and forward rendering to real-world scenarios is bottlenecked by the limited realism and temporal coherence of existing synthetic datasets. To bridge this persistent domain gap, we introduce a large-scale, dynamic dataset curated from visually complex AAA games. Using a novel dual-screen stitched capture method, we extracted 4M continuous frames (720p/30 FPS) of synchronized RGB and five G-buffer channels across diverse scenes, visual effects, and environments, including adverse weather and motion-blur variants. This dataset uniquely advances bidirectional rendering: enabling robust in-the-wild geometry and material decomposition, and facilitating high-fidelity G-buffer-guided video generation. Furthermore, to evaluate the real-world performance of inverse rendering without ground truth, we propose a novel VLM-based assessment protocol measuring semantic, spatial, and temporal consistency. Experiments demonstrate that inverse renderers fine-tuned on our data achieve superior cross-dataset generalization and controllable generation, while our VLM evaluation strongly correlates with human judgment. Combined with our toolkit, our forward renderer enables users to edit styles of AAA games from G-buffers using text prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。