用可学习密度控制生成3D高斯,让模型自动聚焦复杂区域。
Generative 3D Gaussians with Learned Density Control

- 将高斯中心建模为可学习密度采样,实现自适应分布
- 单个隐变量支持多分辨率解码,生成质量达顶尖水平
- 提出新编码机制,加速无序隐变量的扩散生成
我们提出密度采样高斯(DeG),一种新型3D表示方法,旨在弥合自适应渲染原语与可扩展生成建模之间的差距。与传统固定体素网格或数组不同,DeG将高斯中心建模为定义在八叉树上的可学习概率密度函数的采样点。该形式提供了严格的数学框架以实现自适应密度控制:通过在渲染监督下联合优化空间密度与高斯属性,模型自然地将原语集中于几何复杂度高的区域。我们引入一种新的渲染损失梯度,作为标准高斯泼溅中离散致密化和修剪启发式规则的全可微类比。由此产生的表示高度灵活,仅需调整采样预算即可从单一隐变量支持变分辨率解码。为实现生成合成,我们在DeG上训练了隐扩散模型。我们识别出将扩散应用于无序集合结构隐变量时的关键挑战,可能导致收敛显著缓慢,并提出VecSeq——一种将隐令牌锚定到确定性三维Sobol序列的规范重索引机制。这将模糊的集合生成问题转化为稳健的序列建模任务。大量实验表明,我们的流程在单图像到3D生成任务中达到顶尖质量,结合了无结构原语的结构性自适应与基于网格方法的训练稳定性。
原文摘要 · Abstract (English)
We present Density-Sampled Gaussians (DeG), a novel 3D representation designed to bridge the gap between adaptive rendering primitives and scalable generative modeling. Unlike existing approaches that constrain 3D Gaussians to fixed voxel grids or arrays, DeG models Gaussian centers as samples from a learnable probability density function defined over an octree. This formulation provides a rigorous mathematical framework for adaptive density control: by jointly optimizing the spatial density and Gaussian attributes under rendering supervision, our model naturally concentrates primitives in regions of high geometric complexity. We achieve this via a new render loss contribution gradient that serves as a fully differentiable analogue to the discrete densification and pruning heuristics used in standard Gaussian Splatting. The resulting representation is highly flexible, supporting variable-resolution decoding from a single latent code by simply adjusting the sampling budget. To enable generative synthesis, we train a latent diffusion model on DeG. We identify a critical challenge in applying diffusion to unordered set-structured latents, which can significantly slow convergence, and propose VecSeq, a canonical re-indexing mechanism that anchors latent tokens to a deterministic 3D Sobol sequence. This transforms the ambiguous set-generation problem into a robust sequence modeling task. Extensive experiments demonstrate that our pipeline achieves state-of-the-art quality in single-image-to-3D generation, combining the structural adaptivity of unstructured primitives with the training stability of grid-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。