arXiv:2604.12508cs.CV2026-04ACL

解决大模型看细节能力差的问题,通过信息流调控提升视觉感知精度

From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception

论文配图:From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception
图 1 · 摘自论文原文
  • 用变分自编码器建模视觉注意力,动态调节信息流动
  • 在多个细粒度任务上显著提升准确率,尤其在微小目标识别中表现突出
  • 可插拔设计,适合希望增强模型细节理解能力的研究者

尽管多模态大语言模型在通用视觉理解上表现优异,但在需要识别微小物体或辨别细微视觉关系的细粒度感知任务中常表现不佳。我们将其归因于视觉衰减现象:稀疏的细粒度视觉信号在传播过程中被主导的文本令牌过早抑制或稀释,导致深层决策时‘注意力丢失’。现有以输入为中心的解决方案无法从根本上逆转这一内在的信息损耗机制。为此,我们提出变分信息流(VIF)框架,采用概率视角,利用条件变分自编码器(CVAE)将与问答对相关的视觉显著性建模为潜在分布。作为即插即用模块,VIF可集成至现有架构。在涵盖通用VQA、细粒度感知和视觉定位的多个基准上的广泛评估表明,VIF相较先前方法取得具有竞争力的性能提升,验证了其在增强多模态大模型细粒度感知能力方面的有效性。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding, they frequently falter in fine-grained perception tasks that require identifying tiny objects or discerning subtle visual relationships. We attribute this limitation to Visual Attenuation: a phenomenon where sparse fine-grained visual signals are prematurely suppressed or diluted by dominant textual tokens during network propagation, resulting in a "loss of focus" during the deep-level decision-making process. Existing input-centric solutions fail to fundamentally reverse this intrinsic mechanism of information loss. To address this challenge, we propose the Variational Information Flow (VIF) framework. Adopting a probabilistic perspective, VIF leverages a Conditional Variational Autoencoder (CVAE) to model the visual saliency relevant to the question-answer pair as a latent distribution. As a plug-and-play module, VIF can be integrated into existing architectures. Extensive evaluations across diverse benchmarks, covering General VQA, fine-grained perception, and visual grounding, demonstrate that VIF yields competitive improvements over previous methods, validating its effectiveness in enhancing the fine-grained perception of MLLMs.

细粒度感知多模态模型信息流调控视觉注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。