arXiv:2503.06784cs.GRcs.AI2025-03被引 1

用真实海底数据训练,生成逼真3D海底场景。

Infinite Leagues Under the Sea: Photorealistic 3D Underwater Terrain Generation by Latent Fractal Diffusion Models

  • 基于分形潜空间嵌入,结合视觉基础模型提取3D几何与语义。
  • 在真实海底图像基础上生成高保真RGBD图像,支持大尺度一致性渲染。
  • 适合影视、游戏和机器人仿真,可实现逼真视角切换。

本文针对3D海底地形生成问题,提出DreamSea模型。现有通用生成模型因缺乏专业海底图像训练,生成效果失真。DreamSea基于水下机器人调查收集的真实图像数据库进行训练,该数据包含大量真实海底观测且覆盖广阔区域,但存在噪声与现实世界伪影。我们利用视觉基础模型从数据中提取3D几何与语义信息,训练一种扩散模型,通过新型分形分布驱动的潜空间嵌入,生成具有真实感的海底RGBD图像。随后将生成图像融合为3D地图,并构建由2D扩散先验监督的3DGS模型,实现高保真新视角渲染。实验表明,DreamSea能鲁棒生成大规模、一致且多样化的逼真海底场景,在影视、游戏及机器人仿真等领域具广泛应用潜力。

原文摘要 · Abstract (English)

This paper tackles the problem of generating representations of underwater 3D terrain. Off-the-shelf generative models, trained on Internet-scale data but not on specialized underwater images, exhibit downgraded realism, as images of the seafloor are relatively uncommon. To this end, we introduce DreamSea, a generative model to generate hyper-realistic underwater scenes. DreamSea is trained on real-world image databases collected from underwater robot surveys. Images from these surveys contain massive real seafloor observations and covering large areas, but are prone to noise and artifacts from the real world. We extract 3D geometry and semantics from the data with visual foundation models, and train a diffusion model that generates realistic seafloor images in RGBD channels, conditioned on novel fractal distribution-based latent embeddings. We then fuse the generated images into a 3D map, building a 3DGS model supervised by 2D diffusion priors which allows photorealistic novel view rendering. DreamSea is rigorously evaluated, demonstrating the ability to robustly generate large-scale underwater scenes that are consistent, diverse, and photorealistic. Our work drives impact in multiple domains, spanning filming, gaming, and robot simulation.

3D生成海底建模扩散模型虚拟仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。