无需标定与深度,实现跨模态视角合成。
No Calibration, No Depth, No Problem: Cross-Sensor View Synthesis with 3D Consistency
- 提出匹配-加密-整合方法,自动对齐不同传感器图像。
- 在无3D先验下实现高质量视图合成,精度优于现有方法。
- 适合大规模真实场景的多传感器数据构建,降低工程成本。
我们首次研究了跨模态的跨传感器视图合成问题。针对一个实际且被广泛忽视的关键挑战——获取对齐的RGB-X数据:多数现有工作假设此类配对已存在,专注于模态融合,但实际需大量标定工程。为此,我们提出一种匹配-加密-整合方法:先进行RGB-X图像匹配,再通过引导点加密;利用置信度感知加密和自匹配过滤,提升合成质量,并在3D高斯泼溅(3DGS)中进行三维一致性整合。该方法无需X传感器的3D先验,仅依赖近零成本的COLMAP用于RGB。目标是消除多种RGB-X传感器的繁琐标定流程,通过可扩展方案突破大规模真实数据采集瓶颈,推动跨传感器学习普及。
原文摘要 · Abstract (English)
We present the first study of cross-sensor view synthesis across different modalities. We examine a practical, fundamental, yet widely overlooked problem: getting aligned RGB-X data, where most RGB-X prior work assumes such pairs exist and focuses on modality fusion, but it empirically requires huge engineering effort in calibration. We propose a match-densify-consolidate method. First, we perform RGB-X image matching followed by guided point densification. Using the proposed confidence-aware densification and self-matching filtering, we attain better view synthesis and later consolidate them in 3D Gaussian Splatting (3DGS). Our method uses no 3D priors for X-sensor and only assumes nearly no-cost COLMAP for RGB. We aim to remove the cumbersome calibration for various RGB-X sensors and advance the popularity of cross-sensor learning by a scalable solution that breaks through the bottleneck in large-scale real-world RGB-X data collection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。