纯数据驱动的视觉变压器实现无偏立体匹配,精度与速度双突破。
NBS: No Bias Stereo

- 用纯Vision Transformer替代传统几何先验,端到端训练。
- 在大规模合成数据上超越现有方法,精度达最新水平,推理更快。
- 打破对显式先验的依赖,适合追求可扩展3D重建的研究者。
立体重建是计算机视觉中最后一个仍依赖复杂架构先验的任务。尽管已有通用方法能解决该问题,但普遍认为显式几何先验对高质量结果和计算效率至关重要。本文挑战这一范式:我们证明,完全摒弃架构先验的模型也能达到顶尖精度并具备更优运行效率。通过在大规模合成数据上训练一个简单的端到端视觉变压器,我们展示纯数据驱动学习可超越显式设计的几何结构。本工作证实,显式先验不再是立体匹配的必要条件,为3D重建的持续改进打开了真正的可扩展性空间。
原文摘要 · Abstract (English)
Stereo reconstruction is one of the last remaining Computer Vision tasks where all state-of-the-art methods employ a heavy architectural inductive bias. Even though it has been demonstrated that the task can be solved using general-purpose methods, it is widely believed that inductive biases in stereo are strictly necessary for both high-quality results and computational efficiency. We challenge this paradigm. In this paper, we demonstrate that both state-of-the-art accuracy and superior runtime efficiency are achievable with a model completely devoid of architectural inductive biases, relying instead on a simple, end-to-end Vision Transformer. By training on massive synthetic datasets, we show that pure data-driven learning can surpass explicitly engineered geometry. This work proves that explicit inductive biases are no longer a prerequisite for stereo matching, ultimately unlocking true scaling laws for continuous improvement in 3D reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。