arXiv:2511.16555cs.CV2025-11被引 4

轻量模型实现强零样本立体匹配,效率超群。

Lite Any Stereo: Efficient Zero-Shot Stereo Matching

  • 设计紧凑高效的骨干网络与混合代价聚合模块
  • 在百万级数据训练下实现跨场景零样本泛化,四大真实场景基准第一
  • 计算成本不足1%,性能媲美甚至超越主流高精度方法

近期立体匹配研究多聚焦于精度提升,但往往伴随模型规模显著增大。传统观点认为高效模型因容量有限,难以具备零样本能力。本文提出 Lite Any Stereo,一种兼具强零样本泛化能力与高效率的立体深度估计框架。通过设计紧凑而富有表现力的主干网络及精心设计的混合代价聚合模块,并在百万级数据上采用三阶段训练策略,有效弥合了模拟到真实场景的差距。实验表明,该超轻量模型在四个广泛使用的现实世界基准上排名第一。令人瞩目的是,其精度可比肩或超越无需先验的先进非精确方法,同时计算成本低于1%,为高效立体匹配树立了新标准。

原文摘要 · Abstract (English)

Recent advances in stereo matching have focused on accuracy, often at the cost of significantly increased model size. Traditionally, the community has regarded efficient models as incapable of zero-shot ability due to their limited capacity. In this paper, we introduce Lite Any Stereo, a stereo depth estimation framework that achieves strong zero-shot generalization while remaining highly efficient. To this end, we design a compact yet expressive backbone to ensure scalability, along with a carefully crafted hybrid cost aggregation module. We further propose a three-stage training strategy on million-scale data to effectively bridge the sim-to-real gap. Together, these components demonstrate that an ultra-light model can deliver strong generalization, ranking 1st across four widely used real-world benchmarks. Remarkably, our model attains accuracy comparable to or exceeding state-of-the-art non-prior-based accurate methods while requiring less than 1% computational cost, setting a new standard for efficient stereo matching.

立体匹配轻量模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。