两视图训练可显著提升多视角3D高斯点云的重建质量。
Two-View Accumulation as the Primary Training Lever for Hybrid-Capture Gaussian Splatting: A Variance-Decomposition View of When Gradient Surgery Helps

- 每步优化渲染两个视角,简单有效提升性能。
- 在五个基准上平均提升1-3 dB,超越其他复杂方法。
- 适合研究多源视角融合与训练策略的开发者。
混合捕获的新视角合成结合了远距离(如航拍)与近距离(如地面)图像。标准3D高斯点云(3DGS)在五项混合捕获基准上以每步一个视角、训练30K轮次时,对少数视角的拟合误差达1-3 dB。本文识别出关键改进杠杆:在计算量相当的前提下,仅将每步渲染视角数增至两个,即可显著缩小差距。相比梯度归一化、方向感知梯度手术、投影预条件、置信度门控样本级手术及主动损失差异配对等方法,简单增加视图数量效果最优。我们提出方差分解框架解释该现象:在双模态相机分布下,跨域梯度方差远小于域内方差,因此结构化或随机配对无差异,而双视图累积带来的方差减半是主导因素。该框架在五组场景中验证,其相机高度双模性系数范围为[0.55, 1.00]。结果表明,前述复杂方法均未超越随机双视图配对,且该结构优势可迁移至Scaffold-GS与Pixel-GS骨干网络。
原文摘要 · Abstract (English)
Hybrid-capture novel view synthesis combines images at substantially different camera distances (e.g., aerial drone and ground-level views). Standard 3D Gaussian Splatting (3DGS), trained for 30K iterations with one rendered view per optimizer step, under-fits the minority regime by 1-3 dB on five hybrid-capture benchmarks. We isolate the lever that closes this gap. Among compute-matched alternatives -- vanilla 60K iterations, magnitude corrections (GradNorm), direction-aware near/far gradient surgery, projective preconditioning, confidence-gated sample-level surgery, and a random two-view-per-step control -- the simplest structural change wins: rendering two views per optimizer step. The pairing rule (geometry-defined near/far, random, or active loss-disparity) does not change PSNR beyond seed variance on any of the five scenes; the structural change of having two views per step does. We propose a variance-decomposition framework that predicts and explains this finding: under bimodal camera regimes, between-regime gradient variance turns out to be small relative to within-regime variance in 3DGS, so structured and random pairings are variance-equivalent in expectation, and the variance halving from two-view accumulation itself is the dominant effect. We verify the framework on five scenes whose camera-altitude bimodality coefficients span [0.55, 1.00], and we report the negative result that direction-aware projection, magnitude correction, confidence gating, and an active loss-disparity pairing all fall within seed variance of random two-view pairing. The two-view structural lever transfers cleanly to the Scaffold-GS and Pixel-GS backbones. We position this work as an honest characterization of which training-side axes do and do not move PSNR for hybrid-capture 3DGS, together with the framework that explains why.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。