通过频谱分析发现:3D重建质量取决于特征一致性,而非细节锐度。
Spectral Probing of Feature Upsamplers in 2D-to-3D Scene Reconstruction
- 提出六项频谱指标,诊断特征上采样对3D感知的影响
- 结构频谱一致性越强,新视图合成效果越好,高频增强反致性能下降
- 适合研究2D到3D重建的算法设计者,尤其关注特征保真度的场景
典型的2D到3D重建流程以多视角图像为输入,由视觉基础模型(VFM)提取特征,并通过空间上采样生成稠密表示用于3D重建。若不同视角的稠密特征保持几何一致性,可微渲染即可恢复高精度3D结构,因此特征上采样器是关键组件。现有可学习上采样方法主要聚焦提升空间细节(如更锐利的几何或更丰富的纹理),但其对3D感知的影响仍不明确。为此,本文提出一套包含六项互补指标的频谱诊断框架,用于刻画振幅重分布、结构频谱对齐与方向稳定性。在CLIP与DINO骨干网络上对比经典插值与可学习上采样方法,得出三个关键发现:第一,结构频谱一致性(SSC/CSC)是新视图合成(NVS)质量最强预测因子,而高频谱斜率漂移(HFSS)常与重建性能负相关,表明单纯强调高频细节未必提升3D重建;第二,几何与纹理对不同频谱特性响应不同:角能量一致性(ADC)与几何指标更强相关,而SSC/CSC对纹理保真度影响略大于几何准确性;第三,尽管可学习上采样通常生成更锐利的空间特征,却极少优于经典插值法的重建质量,且其有效性依赖于重建模型。总体而言,重建质量更依赖于频谱结构的保留而非空间细节增强,凸显频谱一致性在2D到3D上采样策略设计中的核心作用。
原文摘要 · Abstract (English)
A typical 2D-to-3D pipeline takes multi-view images as input, where a Vision Foundation Model (VFM) extracts features that are spatially upsampled to dense representations for 3D reconstruction. If dense features across views preserve geometric consistency, differentiable rendering can recover an accurate 3D representation, making the feature upsampler a critical component. Recent learnable upsampling methods mainly aim to enhance spatial details, such as sharper geometry or richer textures, yet their impact on 3D awareness remains underexplored. To address this gap, we introduce a spectral diagnostic framework with six complementary metrics that characterize amplitude redistribution, structural spectral alignment, and directional stability. Across classical interpolation and learnable upsampling methods on CLIP and DINO backbones, we observe three key findings. First, structural spectral consistency (SSC/CSC) is the strongest predictor of NVS quality, whereas High-Frequency Spectral Slope Drift (HFSS) often correlates negatively with reconstruction performance, indicating that emphasizing high-frequency details alone does not necessarily improve 3D reconstruction. Second, geometry and texture respond to different spectral properties: Angular Energy Consistency (ADC) correlates more strongly with geometry-related metrics, while SSC/CSC influence texture fidelity slightly more than geometric accuracy. Third, although learnable upsamplers often produce sharper spatial features, they rarely outperform classical interpolation in reconstruction quality, and their effectiveness depends on the reconstruction model. Overall, our results indicate that reconstruction quality is more closely related to preserving spectral structure than to enhancing spatial detail, highlighting spectral consistency as an important principle for designing upsampling strategies in 2D-to-3D pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。