用单目深度引导的高斯点云,让地面机器人场景渲染更准更稳。
Mode-GS: Monocular Depth Guided Anchored 3D Gaussian Splatting for Robust Ground-View Scene Rendering
- 基于单目深度生成像素对齐锚点,构建抗漂移的高斯点云
- 在R3LIVE和Tanks and Temples数据集上实现最佳渲染效果
- 解决单目深度尺度模糊问题,适合无固定轨迹的地面场景
我们提出一种新型视角渲染算法Mode-GS,适用于地面机器人轨迹数据集。该方法采用锚定高斯点云,以克服现有3D高斯溅射算法的局限性。先前的神经渲染方法因场景复杂且视图观测不足,导致点云严重漂移,难以在地面机器人数据集中准确固定于真实几何结构。我们的方法利用单目深度生成像素对齐的锚点,并通过残差形式的高斯解码器在锚点周围生成高斯点云。为解决单目深度固有的尺度模糊问题,我们采用每视图深度尺度参数化,并引入尺度一致的深度损失实现在线尺度校准。实验表明,该方法在具有自由轨迹模式的地面场景中,基于PSNR、SSIM和LPIPS指标均取得更优渲染性能,在R3LIVE位姿估计数据集和Tanks and Temples数据集上达到当前最优水平。
原文摘要 · Abstract (English)
We present a novel-view rendering algorithm, Mode-GS, for ground-robot trajectory datasets. Our approach is based on using anchored Gaussian splats, which are designed to overcome the limitations of existing 3D Gaussian splatting algorithms. Prior neural rendering methods suffer from severe splat drift due to scene complexity and insufficient multi-view observation, and can fail to fix splats on the true geometry in ground-robot datasets. Our method integrates pixel-aligned anchors from monocular depths and generates Gaussian splats around these anchors using residual-form Gaussian decoders. To address the inherent scale ambiguity of monocular depth, we parameterize anchors with per-view depth-scales and employ scale-consistent depth loss for online scale calibration. Our method results in improved rendering performance, based on PSNR, SSIM, and LPIPS metrics, in ground scenes with free trajectory patterns, and achieves state-of-the-art rendering performance on the R3LIVE odometry dataset and the Tanks and Temples dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。