arXiv:2606.21300cs.CV2026-06International Conf…

单帧处理长视频,实现3D几何的精准与一致估计

SCOPE: Scale-Consistent One-Pass Estimation of 3D Geometry

论文配图:SCOPE: Scale-Consistent One-Pass Estimation of 3D Geometry
图 1 · 摘自论文原文
  • 用统一参数生成仿射不变的3D点图,保持跨帧一致性
  • 在ScanNet上点图误差降24.2%,时间对齐误差降34.9%
  • 适合长序列、复杂运动与光照变化场景的3D重建

我们提出SCOPE(Scale-Consistent One-Pass Estimation of 3D Geometry),一种从长单目视频序列中估计3D几何的新方法。现有方法在数百帧的长序列中难以同时保证几何精度与时间一致性。本方法通过共享参数生成仿射不变的3D点图,实现尺度一致的表示。提出三大创新:视点不变的几何对齐,将多视角点映射到统一参考系;外观不变学习,确保跨指数时间尺度的一致性;频调定位,支持外推至远超训练长度的序列。在多个数据集上的实验表明,相比最先进方法,扫描图(ScanNet)上相对点图误差降低24.2%,时间对齐误差降低34.9%。该方法可高效处理复杂相机轨迹与光照变化下的长序列,且仅需单次遍历。

原文摘要 · Abstract (English)

We present SCOPE (Scale-Consistent One-Pass Estimation of 3D Geometry), a novel approach for estimating 3D geometry from extended monocular video sequences, where existing methods struggle to maintain both geometric accuracy and temporal consistency across hundreds of frames. Our approach generates affine-invariant 3D point maps with shared parameters across entire sequences, enabling consistent scale-invariant representations. We introduce three key innovations: viewpoint-invariant geometry aligning multi-perspective points in a unified reference frame; appearance-invariant learning enforcing consistency across exponential timescales; and frequency-modulated positioning enabling extrapolation to sequences vastly exceeding training length. Experiments across diverse datasets demonstrate significant improvements, reducing relative point map error by 24.2% and temporal alignment error by 34.9% on ScanNet compared to state-of-the-art methods. Our approach handles challenging scenarios with complex camera trajectories and lighting variations while efficiently processing extended sequences in a single pass. Project page: https://scope3d.github.io/.

3D重建单目视觉时间一致性长序列处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。