3D高斯溅射评估常误判空间泛化能力,实际测的是近轨迹插值。
Mind the Gap: Standard 3DGS Evaluation Primarily Measures Near-Trajectory Interpolation

- 设计等量配对的隔帧与连续区域两种测试方式,精准区分插值与外推。
- 发现插值与外推性能差距达3~12dB,远超方法间差异。
- 结果适用于多种模型,提示当前评估协议需重构,适合重建与采集规划研究者。
标准的MipNeRF360风格3D高斯溅射(3DGS)评估通常每隔N帧留出一个测试帧——这些帧在训练时两侧均有邻近帧,因此评测实际上衡量的是近轨迹插值能力,而非真正的空间泛化。本文提出一种公平的匹配计数协议:两组训练图像数量相同,仅测试集分布不同——一组均匀分布(插值),另一组为连续空间区域(外推)。主要发现是插值与外推之间存在3~12dB的巨大且一致的性能差距,远超多数方法间的差异。该差距对训练噪声鲁棒,在两个案例中足以改变方法排名,且在三种表示体系(包括非高斯体素神经辐射场)中均持续存在,说明其反映的是空间覆盖本质,而非特定模型特性。诊断分析显示,该差距主要由漫反射/几何代理成分主导,并与每视角到最近训练视角的角距离高度相关,这一零成本信号还可用于指导数据采集规划;损失侧正则化仅带来微弱提升。标准留出策略仍适用于近轨迹渲染,但不应单独作为空间泛化的证据。本文首次结合匹配计数配对留出、跨表示量化与诊断分析,构建了包含16个场景的空间留出基准工具包,即将开源。
原文摘要 · Abstract (English)
Standard MipNeRF360-style 3D Gaussian Splatting (3DGS) evaluation holds out every N-th frame -- but these frames have trained neighbors on both sides, so the metric measures near-trajectory interpolation rather than spatial generalization. We introduce a fair matched-count protocol that isolates this effect: both arms train on the same number of images and differ only in whether the holdout is spread evenly (interpolation) or forms a contiguous spatial sector (extrapolation). Our primary finding is a large, consistent interpolation-extrapolation gap of 3~12dB -- several times the differences typically reported between competing methods. The gap is robust to training noise, is in two cases large enough to flip a method ranking under multi-seed confirmation, and -- crucially -- persists across three representation families, including a non-Gaussian volumetric neural radiance field (NeRF), so it reflects spatial coverage rather than any one representation. Diagnostically, it is dominated by a diffuse/geometry-proxy component and tracks each view's angular distance to its nearest training view, a zero-cost signal that also guides capture planning; loss-side regularization yields only marginal gains. Standard holdouts remain useful for near-trajectory rendering but should not, alone, be read as evidence of spatial generalization. Prior work notes protocol sensitivity; ours is, to our knowledge, the first to combine matched-count paired holdout, cross-representation quantification, and a diagnostic analysis Table 1. We describe a spatial-holdout benchmark toolkit with standardized splits and baselines for 16 scenes, which we are preparing for public release.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。