arXiv:2608.28933cs.CV2026-08

纯数据驱动的视觉变压器实现无偏立体匹配,精度与速度双突破。

NBS: No Bias Stereo

论文配图:NBS: No Bias Stereo
图 1 · 摘自论文原文
  • 用纯Vision Transformer替代传统几何先验,端到端训练。
  • 在大规模合成数据上超越现有方法,精度达最新水平,推理更快。
  • 打破对显式先验的依赖,适合追求可扩展3D重建的研究者。

立体重建是计算机视觉中最后一个仍依赖复杂架构先验的任务。尽管已有通用方法能解决该问题,但普遍认为显式几何先验对高质量结果和计算效率至关重要。本文挑战这一范式:我们证明,完全摒弃架构先验的模型也能达到顶尖精度并具备更优运行效率。通过在大规模合成数据上训练一个简单的端到端视觉变压器,我们展示纯数据驱动学习可超越显式设计的几何结构。本工作证实,显式先验不再是立体匹配的必要条件,为3D重建的持续改进打开了真正的可扩展性空间。

原文摘要 · Abstract (English)

Stereo reconstruction is one of the last remaining Computer Vision tasks where all state-of-the-art methods employ a heavy architectural inductive bias. Even though it has been demonstrated that the task can be solved using general-purpose methods, it is widely believed that inductive biases in stereo are strictly necessary for both high-quality results and computational efficiency. We challenge this paradigm. In this paper, we demonstrate that both state-of-the-art accuracy and superior runtime efficiency are achievable with a model completely devoid of architectural inductive biases, relying instead on a simple, end-to-end Vision Transformer. By training on massive synthetic datasets, we show that pure data-driven learning can surpass explicitly engineered geometry. This work proves that explicit inductive biases are no longer a prerequisite for stereo matching, ultimately unlocking true scaling laws for continuous improvement in 3D reconstruction.

立体匹配视觉变压器无偏学习3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。