arXiv:2608.23206cs.CV2026-08

用球形占据轮廓统一建模多视角3D重建与生成,兼顾精度与可控性。

Learning Spherical Occupancy Profiles for Multi-View 3D Reconstruction and Generation

论文配图:Learning Spherical Occupancy Profiles for Multi-View 3D Reconstruction and Generation
图 1 · 摘自论文原文
  • 将多视角高斯重建提炼为射线级占据概率轮廓,作为统一中间表示
  • 判别式模型达0.035(归一化)中位深度误差,生成模型支持可调解空间多样性
  • 轮廓形态可优化,适合需要不确定性感知的3D生成与重建任务

我们研究球形占据轮廓——从多视角3D高斯重建中提炼出的射线级占据概率分布P(r) = T(r) o(r),作为判别与生成式3D重建的统一中间表示。在包含999个物体、每物体48个旋转视角的Google Scanned Objects子集上,训练了:(i) 一个判别式逐射线解码器,通过FiLM条件化的轮廓头注入全局视图平均与射线特定图像证据,在独立90个物体测试集上达到0.035(归一化)中位软深度误差;(ii) 一个基于轮廓变分自编码器(VAE)与潜在扩散模型的生成流程,支持无条件采样以匹配重建流形,以及图像条件下的多解重建,其每物体解空间分布可通过无分类器引导定量调控。进一步分析表明,后处理幂锐化与学习的锐化目标均可恢复真实轮廓宽度而不降低深度精度,揭示了L1逐射线损失族中宽度-峰值单调前沿,启发对形态门控的合理重定义。在两个DTU真实场景上的照片验证确认该流程可迁移到非合成输入。结果表明,射线级占据轮廓提供了紧凑、可学习且带不确定性的多视角重建与生成先验接口。

原文摘要 · Abstract (English)

We study spherical occupancy profiles-the ray-wise occupancy probability profiles P(r) = T(r) o(r) distilled from multi-view 3D Gaussian reconstructions-as a unified intermediate representation for both discriminative and generative 3D reconstruction from images. On a 999-object subset of Google Scanned Objects with 48 turntable views each, we train (i) a discriminative per-ray decoder that injects global view-averaged and ray-specific image evidence into a FiLM-conditioned profile head, reaching median soft depth error 0.035 (normalized) on an independent 90-object test split, and (ii) a generative pipeline built on a profile VAE and a latent diffusion model, which supports unconditional sampling that matches the reconstruction manifold and image-conditioned multi-solution reconstruction whose per-object solution spread is quantifiable and tunable via classifier-free guidance. We further analyze the morphology of predicted profiles: post-hoc power sharpening and a learned sharpening target both recover ground-truth profile width without degrading depth, exposing a monotonic width-peak frontier in the L1-per-ray loss family and motivating a principled redefinition of morphology gates. Real-photo validation on two DTU scenes confirms the pipeline transfers to non-synthetic input. Our results suggest that ray-wise occupancy profiles offer a compact, learned, and uncertainty-aware interface between multi-view reconstruction and generative priors.

3D重建生成模型高斯渲染轮廓表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。