arXiv:2604.13905cs.CV2026-04

用稀疏查询实现高效3D生成,减少视角偏差。

Rethinking Image-to-3D Generation with Sparse Queries: Efficiency, Capacity, and Input-View Bias

论文配图:Rethinking Image-to-3D Generation with Sparse Queries: Efficiency, Capacity, and Input-View Bias
图 1 · 摘自论文原文
  • 用可学习的稀疏3D锚点加扩展算子生成3D高斯点。
  • 内存和推理时间大幅降低,多视角一致性保持良好。
  • 适合追求高效且低视角依赖的3D生成研究者。

我们提出SparseGen,一种高效的图像到3D生成框架,显著降低输入视角偏差。不同于依赖密集体素网格、三平面或像素对齐基元的传统方法,该框架采用紧凑的稀疏学习3D锚点集与学习的扩展算子,将每个变换后的锚点解码为少量局部3D高斯基元。在无3D监督的修正流重建目标下训练,模型能将表示容量分配至几何与外观关键区域,实现内存与推理时间的显著减少,同时保持多视角保真度。我们引入输入视角偏差与利用率的量化度量,证明稀疏查询能降低对条件视图的过拟合,同时具备高表示效率。结果表明,稀疏集合-隐式扩展是一种原理清晰且实用的高效3D生成建模替代方案。

原文摘要 · Abstract (English)

We present SparseGen, a novel framework for efficient image-to-3D generation, which exhibits low input-view bias while being significantly faster. Unlike traditional approaches that rely on dense volumetric grids, triplanes, or pixel-aligned primitives, we model scenes with a compact sparse set of learned 3D anchor queries and a learned expansion operator that decodes each transformed query into a small local set of 3D Gaussian primitives. Trained under a rectified-flow reconstruction objective without 3D supervision, our model learns to allocate representation capacity where geometry and appearance matter, achieving significant reductions in memory and inference time while preserving multi-view fidelity. We introduce quantitative measures of input-view bias and utilization to show that sparse queries reduce overfitting to conditioning views while being representationally efficient. Our results argue that sparse set-latent expansion is a principled, practical alternative for efficient 3D generative modeling.

3D生成稀疏表示高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。