arXiv:2512.10548cs.CV2025-12被引 2

让视觉模型像人一样动态聚焦关键区域,提升理解能力。

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

  • 模拟人类眨眼式扫描,动态调整视觉令牌关注
  • 在多个数据集上显著提升视觉理解准确率
  • 适合需要高效精准视觉感知的多模态任务

多模态大语言模型在视觉-语言任务中取得显著进展,但其视觉感知能力仍有限。人类通过逐次“眨眼式”扫描和聚焦显著区域来高效感知复杂场景。受此启发,我们首次探究多模态大模型是否具备类似行为。初步分析发现,多模态大模型在不同层自然关注不同视觉区域,且对显著令牌分配更多计算可增强视觉感知。基于此,我们提出Blink:一种在单次前向传播中模拟人类启发式过程的动态视觉令牌分辨率框架。该框架包含两个模块:显著性引导扫描与动态令牌分辨率。它首先根据注意力图估计各层视觉令牌的显著性,并通过即插即用的令牌超分辨率(TokenSR)模块扩展重要令牌。在下一层中,当令牌失去关注时将其丢弃。这种动态机制在广泛探索与精细聚焦间取得平衡,从而自适应、高效地提升视觉感知。大量实验证明Blink能有效增强视觉感知与多模态理解能力。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have achieved remarkable progress on various vision-language tasks, yet their visual perception remains limited. Humans, in comparison, perceive complex scenes efficiently by dynamically scanning and focusing on salient regions in a sequential "blink-like" process. Motivated by this strategy, we first investigate whether MLLMs exhibit similar behavior. Our pilot analysis reveals that MLLMs naturally attend to different visual regions across layers and that selectively allocating more computation to salient tokens can enhance visual perception. Building on this insight, we propose Blink, a dynamic visual token resolution framework that emulates the human-inspired process within a single forward pass. Specifically, Blink includes two modules: saliency-guided scanning and dynamic token resolution. It first estimates the saliency of visual tokens in each layer based on the attention map, and extends important tokens through a plug-and-play token super-resolution (TokenSR) module. In the next layer, it drops the extended tokens when they lose focus. This dynamic mechanism balances broad exploration and fine-grained focus, thereby enhancing visual perception adaptively and efficiently. Extensive experiments validate Blink, demonstrating its effectiveness in enhancing visual perception and multimodal understanding.

多模态视觉感知动态聚焦注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。