arXiv:2511.11177cs.CV2025-11被引 1

用状态空间模型替代注意力,实现快速精准的多模态理解

Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation

  • 用液态状态空间动态替代传统注意力,降低计算复杂度
  • 在多个基准上实现细粒度视觉理解,推理速度显著提升
  • 适合资源受限场景如机器人、智能摄像头等实际应用

多模态大语言模型在视觉语言理解上取得显著进展,但高计算成本限制了其在机器人操作、个人助手和智能摄像头等资源受限场景中的部署。现有方法多依赖二次复杂度的Transformer交叉注意力,效率低下。小型视觉语言模型常难以精确捕捉细粒度任务相关视觉区域,影响细粒度推理性能。为此,我们提出Viper-F1,一种混合状态空间视觉语言模型,用高效的液态状态空间动态替代注意力机制。为进一步增强视觉定位能力,我们设计了文本标记与图像网格的相关性模块,通过FiLM条件调节状态空间动态,实现对文本提示相关的视觉区域的选择性强调,同时保持线性时间推理。在多个基准上的实验表明,Viper-F1在显著提升效率的同时实现了准确的细粒度理解。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have enabled impressive progress in vision-language understanding, yet their high computational cost limits deployment in resource-constrained scenarios such as robotic manipulation, personal assistants, and smart cameras. Most existing methods rely on Transformer-based cross-attention, whose quadratic complexity hinders efficiency. Moreover, small vision-language models often struggle to precisely capture fine-grained, task-relevant visual regions, leading to degraded performance on fine-grained reasoning tasks that limit their effectiveness in the real world. To address these issues, we introduce Viper-F1, a Hybrid State-Space Vision-Language Model that replaces attention with efficient Liquid State-Space Dynamics. To further enhance visual grounding, we propose a Token-Grid Correlation Module, which computes lightweight correlations between text tokens and image patches and modulates the state-space dynamics via FiLM conditioning. This enables the model to selectively emphasize visual regions relevant to the textual prompt while maintaining linear-time inference. Experimental results across multiple benchmarks demonstrate that Viper-F1 achieves accurate, fine-grained understanding with significantly improved efficiency.

多模态状态空间高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。