arXiv:2608.29819cs.CV2026-08

用频域相位信息提升立体匹配精度,实现实时高精度效果

PhasorNet: Learning Structure from Frequency for Real-Time Stereo Matching

论文配图:PhasorNet: Learning Structure from Frequency for Real-Time Stereo Matching
图 1 · 摘自论文原文
  • 引入傅里叶相位信息增强注意力机制,强化结构一致性
  • 在ETH3D上达顶尖性能,参数仅5.3M,KITT上跨域泛化优异
  • 适合需要实时、高精度立体匹配的自动驾驶与机器人场景

准确的立体匹配在细结构、反光或透明物体等病态区域仍具挑战性,因外观线索常模糊或不可靠。为此,我们提出PhasorNet,一种轻量但强大的框架,通过频域线索提升几何判别力。核心是相位增强变压器(PAT),将傅里叶导出的相位信息注入注意力机制,生成对光照鲁棒、保留结构的特征,优先关注困难区域的结构一致性。此外,我们设计了几何上下文融合精炼模块(GCFRM),结合全分辨率卷积流与轻量注意力流(使用WQA和CDGA模块),高效保持细节与边界,开销小。训练通过多尺度边缘引导高误差区域(EHR)损失增强,自适应聚焦于高误差与边缘区域,指导层级成本体精炼。仅530万参数,PhasorNet在挑战性ETH3D基准上达到领先性能,并在KITTI上展现优秀跨域泛化能力,为高精度实时立体匹配提供高效实用方案。

原文摘要 · Abstract (English)

Accurate stereo matching remains challenging in ill-posed regions such as fine structures, reflective, or transparent objects, where appearance cues are often ambiguous or unreliable. To tackle this, we propose PhasorNet, a lightweight yet powerful framework that boosts geometric discrimination via frequency-domain cues. At its core, the Phase-Augmented Transformer (PAT) injects Fourier-derived phase information into the attention mechanism, yielding photometrically robust, structure-preserving features that prioritize structural consistency in difficult areas. Additionally, we develop a Geometry-Context Fusion Refinement Module (GCFRM) that combines a full-resolution convolutional stream with a lightweight attention-based stream (leveraging WQA and CDGA blocks) to efficiently preserve fine details and object boundaries without excessive overhead. Training is further enhanced by a multi-scale Edge-guided High-Error Region (EHR) loss that adaptively focuses optimization on high-error and edge regions, guiding hierarchical cost volume refinement. With only 5.3M parameters, PhasorNet achieves state-of-the-art performance on the challenging ETH3D benchmark while exhibiting excellent cross-domain generalization on KITTI, delivering an efficient and practical solution for accurate real-time stereo matching.

立体匹配频域分析轻量化模型实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。