arXiv:2503.18544cs.CV2025-03

将知识蒸馏用于立体匹配网络,实现更快更小且性能更强的模型。

Distilling Stereo Networks for Performant and Efficient Leaner Networks

  • 融合先进立体方法与通用蒸馏技术,构建端到端蒸馏框架。
  • 在SceneFlow上超越PSMNet等模型,速度提升5至8倍。
  • 在ETH3D、Middlebury等未见数据集上表现更优,泛化性强。

知识蒸馏在分类和分割任务中已广泛应用,但在先进的立体匹配方法中仍较少研究。原因之一是这些网络结构复杂,通常包含多个二维和三维模块。本文系统结合最新立体匹配方法与通用知识蒸馏技术,提出一种联合蒸馏框架,在保持竞争力的同时实现更快推理。通过详尽的实证分析,我们发现立体匹配网络的蒸馏需精心设计从主干到蒸馏点选择及损失函数的完整流程。结果表明,学生网络不仅更轻量、更快,且性能优异:在SceneFlow基准上,相比PSMNet、CFNet和LEAStereo,性能更优,推理速度分别快8倍、5倍和8倍;相较于推理时间低于100ms的速度优化方法,本模型性能全面超越。此外,在未见过的ETH3D和Middlebury数据集上也展现出更强的泛化能力。

原文摘要 · Abstract (English)

Knowledge distillation has been quite popular in vision for tasks like classification and segmentation however not much work has been done for distilling state-of-the-art stereo matching methods despite their range of applications. One of the reasons for its lack of use in stereo matching networks is due to the inherent complexity of these networks, where a typical network is composed of multiple two- and three-dimensional modules. In this work, we systematically combine the insights from state-of-the-art stereo methods with general knowledge-distillation techniques to develop a joint framework for stereo networks distillation with competitive results and faster inference. Moreover, we show, via a detailed empirical analysis, that distilling knowledge from the stereo network requires careful design of the complete distillation pipeline starting from backbone to the right selection of distillation points and corresponding loss functions. This results in the student networks that are not only leaner and faster but give excellent performance . For instance, our student network while performing better than the performance oriented methods like PSMNet [1], CFNet [2], and LEAStereo [3]) on benchmark SceneFlow dataset, is 8x, 5x, and 8x faster respectively. Furthermore, compared to speed oriented methods having inference time less than 100ms, our student networks perform better than all the tested methods. In addition, our student network also shows better generalization capabilities when tested on unseen datasets like ETH3D and Middlebury.

立体匹配知识蒸馏轻量化高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。