arXiv:2508.01464cs.CV2025-08ICCV被引 5

首个可生成场景级3D高斯的压缩模型,解决尺度不一难题

Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians

  • 提出统一的3D高斯场景编码框架,实现海量点元低维嵌入
  • 在DL3DV-10K数据集上唯一实现新场景泛化,训练失败率低于5%
  • 适用于图像/文本到3D高斯的生成任务,支持下游应用

3D生成虽取得进展,但多局限于物体级别。由于缺乏能扩展至场景级数据的潜在表示学习模型,前馈式3D场景生成仍鲜有探索。与物体级生成不同,以3D高斯点云(3DGS)表示的场景级数据无边界且跨场景尺度不一致,导致统一潜在表示学习极为困难。本文提出首个场景级变分自编码器Can3Tok,可将大量高斯点元编码为低维潜在嵌入,有效捕捉输入的语义与空间信息。此外,我们设计通用数据处理流程以解决尺度不一致问题。在最新场景级3D数据集DL3DV-10K上的验证表明,仅Can3Tok成功泛化至新场景,其他方法在数百个场景输入下即无法收敛,推理时零泛化能力。最后,我们展示了图像到3DGS与文本到3DGS生成的应用,证明其支持下游生成任务的能力。

原文摘要 · Abstract (English)

3D generation has made significant progress, however, it still largely remains at the object-level. Feedforward 3D scene-level generation has been rarely explored due to the lack of models capable of scaling-up latent representation learning on 3D scene-level data. Unlike object-level generative models, which are trained on well-labeled 3D data in a bounded canonical space, scene-level generations with 3D scenes represented by 3D Gaussian Splatting (3DGS) are unbounded and exhibit scale inconsistency across different scenes, making unified latent representation learning for generative purposes extremely challenging. In this paper, we introduce Can3Tok, the first 3D scene-level variational autoencoder (VAE) capable of encoding a large number of Gaussian primitives into a low-dimensional latent embedding, which effectively captures both semantic and spatial information of the inputs. Beyond model design, we propose a general pipeline for 3D scene data processing to address scale inconsistency issue. We validate our method on the recent scene-level 3D dataset DL3DV-10K, where we found that only Can3Tok successfully generalizes to novel 3D scenes, while compared methods fail to converge on even a few hundred scene inputs during training and exhibit zero generalization ability during inference. Finally, we demonstrate image-to-3DGS and text-to-3DGS generation as our applications to demonstrate its ability to facilitate downstream generation tasks.

3D生成高斯点云潜在建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。