arXiv:2606.25437cs.CV2026-06被引 1

用线性全局注意力提升立体匹配,尤其擅长水下等弱光照场景。

LinStereo: Linear-Complexity Global Attention for Multi-Scale Iterative Stereo Matching

论文配图:LinStereo: Linear-Complexity Global Attention for Multi-Scale Iterative Stereo Matching
图 1 · 摘自论文原文
  • 引入位置感知线性注意力,全局聚合信息且计算成本低。
  • 在TartanAir-UW上绝对相对误差降低28%,SQUID上降26%。
  • 适合需要强泛化能力的水下或低光照立体视觉任务。

现有基于视觉基础模型(VFM)的迭代立体匹配流程存在三个短板:多尺度主干特征被压缩为单层相关性、几何先验未在初始化阶段利用、上下文传播仅限局部。这些缺陷在弱光度线索下加剧,使水下场景成为严峻的泛化测试。为此,我们提出LinStereo,基于Depth Anything V3,核心为位置感知线性注意力(PALA)模块,以线性开销替代局部递归,实现全局信息聚合,将可靠估计从匹配良好区域传播至退化区域,同时保持视差结构。PALA的有效性依赖于两个关键组件:层次语义代价体积(HSCV),从VFM特征层级提供对齐尺度的相关性;深度先验初始化(DPI),将单目深度转换为度量校准的初始状态。LinStereo在标准基准上达到顶尖水平,跨域泛化能力强,尤其在严重光度退化的水下场景中表现突出,在TartanAir-UW上绝对相对误差降低28%,在真实水下数据集SQUID上降低26%。

原文摘要 · Abstract (English)

Existing Vision Foundation Model (VFM)-based iterative stereo pipelines under-exploit three information pathways: multi-scale backbone features are collapsed into single-level correlations, geometric priors remain untapped at initialization, and context propagates only locally. These gaps widen under degraded photometric cues, making underwater scenes a stringent generalization test. To address this, we propose LinStereo, built upon Depth Anything V3, whose core is a Position-Aware Linear Attention (PALA) module that replaces local recurrence with global aggregation at linear cost, propagating reliable estimates from well-matched regions into degraded areas while preserving disparity structure. PALA is made effective by two enabling components: Hierarchical Semantic Cost Volumes (HSCV), which supply scale-aligned correlations from the VFM feature hierarchy, and a Depth Prior Initialization (DPI) that converts monocular depth into a metrically calibrated warm start. LinStereo achieves state-of-the-art-level accuracy on standard benchmarks and strong cross-domain generalization, particularly on underwater scene where severe photometric degradation makes stereo matching particularly challenging, attaining the best overall accuracy with consistent gains 28% lower AbsRel on TartanAir-UW, 26% on SQUID, a real-world underwater dataset).

立体匹配水下视觉线性注意力深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。