arXiv:2508.11379cs.CVcs.AI2025-08被引 8

用相机和深度信息引导3D重建,提升精度与灵活性。

G-CUT3R: Guided 3D Reconstruction with Camera and Depth Prior Integration

  • 引入多模态先验编码器,融合深度、相机参数等辅助信息。
  • 在多个基准上实现显著性能提升,最高增益达12.3%。
  • 轻量改造原模型,适配不同输入组合,适合实际场景应用。

我们提出G-CUT3R,一种新型前馈式3D场景重建方法,通过整合先验信息增强CUT3R模型。与仅依赖输入图像的现有前馈方法不同,本方法利用现实中常见的深度图、相机标定或相机位姿等辅助数据。我们对CUT3R进行轻量级改进,为每种模态设计专用编码器提取特征,并通过零卷积将这些特征与RGB图像令牌融合。该灵活设计支持推理时任意组合先验信息。在多个基准(包括3D重建及其他多视角任务)上的评估表明,本方法显著提升性能,证明其能有效利用可用先验,同时兼容不同输入模态。

原文摘要 · Abstract (English)

We introduce G-CUT3R, a novel feed-forward approach for guided 3D scene reconstruction that enhances the CUT3R model by integrating prior information. Unlike existing feed-forward methods that rely solely on input images, our method leverages auxiliary data, such as depth, camera calibrations, or camera positions, commonly available in real-world scenarios. We propose a lightweight modification to CUT3R, incorporating a dedicated encoder for each modality to extract features, which are fused with RGB image tokens via zero convolution. This flexible design enables seamless integration of any combination of prior information during inference. Evaluated across multiple benchmarks, including 3D reconstruction and other multi-view tasks, our approach demonstrates significant performance improvements, showing its ability to effectively utilize available priors while maintaining compatibility with varying input modalities.

3D重建多模态融合视觉几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。