arXiv:2412.16745cs.CV2024-12被引 2

用视觉马尔可夫架构实现快速高精度视差图生成

ViM-Disparity: Bridging the Gap of Speed, Accuracy and Memory for Disparity Map Generation

  • 基于视觉马尔可夫结构,平衡速度、精度与内存消耗
  • 提出联合评估指标,同时衡量推理速度、计算开销和准确率
  • 适合实时立体视觉应用,如自动驾驶与机器人导航

本文提出一种基于视觉马尔可夫(ViM)的架构,旨在解决实时、高精度视差图生成(DMG)模型在计算开销上的固有权衡。我们还设计了一种性能度量方法,可联合评估模型的推理速度、计算开销和准确性。代码与模型已在GitHub公开:https://github.com/MBora/ViM-Disparity。

原文摘要 · Abstract (English)

In this work we propose a Visual Mamba (ViM) based architecture, to dissolve the existing trade-off for real-time and accurate model with low computation overhead for disparity map generation (DMG). Moreover, we proposed a performance measure that can jointly evaluate the inference speed, computation overhead and the accurateness of a DMG model. The code implementation and corresponding models are available at: https://github.com/MBora/ViM-Disparity.

视差图生成视觉马尔可夫实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。