用结构与运动提示提升单目深度模型的立体匹配零样本泛化能力
PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts
- 基于单目深度基础模型的解码器设计新迭代模块PRU
- 在多个数据集上达到顶尖零样本性能,推理速度相当或更快
- 适合关注零样本立体匹配与提示工程的研究者
当前立体匹配方法利用单目深度基础模型实现了优异的零样本泛化性能。然而,多数方法仅聚焦于成本体构建或视差初始化所需的鲁棒特征提取,而对同样关键的迭代优化阶段关注不足。部分方法将单目深度先验作为迭代引导,但传统基于GRU的架构因表征能力有限难以有效利用。本文提出提示循环单元(PRU),一种基于单目深度基础模型解码器的新迭代优化模块。通过将单目结构和立体运动线索作为提示融入解码器,PRU在保留单目深度先验的同时,为模型注入绝对立体尺度信息。实验表明,PromptStereo在多个数据集上实现当前最优的零样本泛化性能,且推理速度相当或更快。研究揭示了提示引导的迭代优化在零样本立体匹配中的巨大潜力。
原文摘要 · Abstract (English)
Modern stereo matching methods have leveraged monocular depth foundation models to achieve superior zero-shot generalization performance. However, most existing methods primarily focus on extracting robust features for cost volume construction or disparity initialization. At the same time, the iterative refinement stage, which is also crucial for zero-shot generalization, remains underexplored. Some methods treat monocular depth priors as guidance for iteration, but conventional GRU-based architectures struggle to exploit them due to the limited representation capacity. In this paper, we propose Prompt Recurrent Unit (PRU), a novel iterative refinement module based on the decoder of monocular depth foundation models. By integrating monocular structure and stereo motion cues as prompts into the decoder, PRU enriches the latent representations of monocular depth foundation models with absolute stereo-scale information while preserving their inherent monocular depth priors. Experiments demonstrate that our PromptStereo achieves state-of-the-art zero-shot generalization performance across multiple datasets, while maintaining comparable or faster inference speed. Our findings highlight prompt-guided iterative refinement as a promising direction for zero-shot stereo matching.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。