用文本生成3D高斯点云,支持一键编辑和多视角合成。
SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis

- 基于多视角校正流生成图像、深度与相机位姿,统一处理文本到3D的转换。
- 在MVImgNet和DL3DV-7K数据集上实现零训练反演与高效3D编辑,精度优于现有方法。
- 适合想快速生成或修改3D场景的设计师与开发者,无需复杂流程。
基于文本的3D场景生成与编辑具有显著潜力,可借助直观用户交互简化内容创作。尽管近期研究利用3D高斯点云(3DGS)实现高保真、实时渲染,但现有方法通常功能专一,缺乏统一的生成与编辑框架。本文提出SplatFlow,一个完整框架,支持直接的3DGS生成与编辑。SplatFlow包含两个核心组件:多视角校正流(RF)模型与高斯点云解码器(GSDecoder)。多视角RF模型在隐空间中同时生成多视角图像、深度图与相机位姿,以文本提示为条件,有效应对真实场景中多样的物体尺度与复杂相机轨迹。随后,GSDecoder通过前馈式3DGS方法将隐空间输出高效转换为3DGS表示。结合无训练反演与图像修复技术,SplatFlow实现无缝3DGS编辑,支持对象编辑、新视角合成与相机位姿估计等多种任务,且无需额外复杂管道。我们在MVImgNet与DL3DV-7K数据集上验证了SplatFlow的能力,展示了其在多种3D生成、编辑与基于修复的任务中的泛化性与有效性。
原文摘要 · Abstract (English)
Text-based generation and editing of 3D scenes hold significant potential for streamlining content creation through intuitive user interactions. While recent advances leverage 3D Gaussian Splatting (3DGS) for high-fidelity and real-time rendering, existing methods are often specialized and task-focused, lacking a unified framework for both generation and editing. In this paper, we introduce SplatFlow, a comprehensive framework that addresses this gap by enabling direct 3DGS generation and editing. SplatFlow comprises two main components: a multi-view rectified flow (RF) model and a Gaussian Splatting Decoder (GSDecoder). The multi-view RF model operates in latent space, generating multi-view images, depths, and camera poses simultaneously, conditioned on text prompts, thus addressing challenges like diverse scene scales and complex camera trajectories in real-world settings. Then, the GSDecoder efficiently translates these latent outputs into 3DGS representations through a feed-forward 3DGS method. Leveraging training-free inversion and inpainting techniques, SplatFlow enables seamless 3DGS editing and supports a broad range of 3D tasks-including object editing, novel view synthesis, and camera pose estimation-within a unified framework without requiring additional complex pipelines. We validate SplatFlow's capabilities on the MVImgNet and DL3DV-7K datasets, demonstrating its versatility and effectiveness in various 3D generation, editing, and inpainting-based tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。