arXiv:2411.08033cs.CVcs.AI2024-11被引 12

用点云潜空间实现3D生成与交互编辑,支持文本、图像和点云输入。

GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation

  • 设计点云结构的潜空间,保留3D形状信息
  • 在多个数据集上优于现有方法,支持多模态条件生成
  • 天然实现几何与纹理解耦,支持3D感知编辑

尽管3D内容生成已取得显著进展,现有方法仍面临输入格式、潜空间设计和输出表示的挑战。本文提出一种新框架GaussianAnything,通过使用多视角姿态对齐的RGB-D-N渲染作为输入,采用独特的点云结构潜空间设计,有效保留3D形状信息,并引入级联潜流模型提升形状-纹理解耦能力。该方法支持多模态条件生成,可接受点云、文本描述或单张图像作为输入。新潜空间天然实现几何与纹理解耦,从而支持3D感知编辑。实验表明,在多个数据集上,该方法在文本和图像条件下的3D生成任务中均优于现有原生3D方法。

原文摘要 · Abstract (English)

While 3D content generation has advanced significantly, existing methods still face challenges with input formats, latent space design, and output representations. This paper introduces a novel 3D generation framework that addresses these challenges, offering scalable, high-quality 3D generation with an interactive Point Cloud-structured Latent space. Our framework employs a Variational Autoencoder (VAE) with multi-view posed RGB-D(epth)-N(ormal) renderings as input, using a unique latent space design that preserves 3D shape information, and incorporates a cascaded latent flow-based model for improved shape-texture disentanglement. The proposed method, GaussianAnything, supports multi-modal conditional 3D generation, allowing for point cloud, caption, and single image inputs. Notably, the newly proposed latent space naturally enables geometry-texture disentanglement, thus allowing 3D-aware editing. Experimental results demonstrate the effectiveness of our approach on multiple datasets, outperforming existing native 3D methods in both text- and image-conditioned 3D generation.

3D生成点云潜空间多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。