arXiv:2412.09912cs.CV2024-12中稿 · AAAI被引 19

用多模型知识提升立体匹配,性能登顶多个榜单

All-in-One: Transferring Vision Foundation Models into Stereo Matching

  • 从多个视觉大模型中选择性迁移知识到立体匹配模型
  • 在Middlebury和ETH3D上分别排名第一,超越已有方法
  • 适合想用大模型提升立体匹配性能的研究者

作为基础视觉任务,立体匹配已取得显著进展。尽管近期基于迭代优化的方法表现优异,其特征提取能力仍有提升空间。受视觉基础模型(VFMs)提取通用表征能力的启发,本文提出AIO-Stereo,可灵活选择并迁移多个异构视觉基础模型的知识至单一立体匹配模型。为更好融合异构模型与立体匹配模型之间的特征,并充分挖掘视觉基础模型中的先验知识,我们设计了双层特征利用机制,实现特征对齐与多层级知识传递。基于此机制,进一步构建双层选择性知识迁移模块,实现知识的选择性转移并整合多个视觉基础模型的优势。实验结果表明,AIO-Stereo在多个数据集上达到顶尖性能,在Middlebury数据集上排名第一,在ETH3D基准上超越所有已发表工作。

原文摘要 · Abstract (English)

As a fundamental vision task, stereo matching has made remarkable progress. While recent iterative optimization-based methods have achieved promising performance, their feature extraction capabilities still have room for improvement. Inspired by the ability of vision foundation models (VFMs) to extract general representations, in this work, we propose AIO-Stereo which can flexibly select and transfer knowledge from multiple heterogeneous VFMs to a single stereo matching model. To better reconcile features between heterogeneous VFMs and the stereo matching model and fully exploit prior knowledge from VFMs, we proposed a dual-level feature utilization mechanism that aligns heterogeneous features and transfers multi-level knowledge. Based on the mechanism, a dual-level selective knowledge transfer module is designed to selectively transfer knowledge and integrate the advantages of multiple VFMs. Experimental results show that AIO-Stereo achieves start-of-the-art performance on multiple datasets and ranks $1^{st}$ on the Middlebury dataset and outperforms all the published work on the ETH3D benchmark.

立体匹配视觉大模型知识迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。