解决多模态大模型看不清细节的结构性盲区问题
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
- 用注意力熵动态生成分层视觉裁剪组合
- 在多个基准上显著超越单一裁剪和无序多裁剪
- 适合需要精准视觉理解的科研与工业场景
多模态大语言模型虽具备强大推理能力,却常忽略细粒度视觉细节,限制其在高精度任务中的应用。现有通过裁剪显著区域的方法存在关键缺陷:'上下文盲区'。该问题源于高保真细节(来自裁剪)与原始图像全局上下文之间的结构断裂,即使所有必要信息都存在。我们指出,这并非信息量不足,而是输入缺乏'结构多样性'。为此,提出无需训练的两步方法Visual Funnel:第一步通过单次前向传播完成上下文锚定,定位关注区域;第二步构建基于注意力熵动态调节的熵缩放组合,自适应确定裁剪尺寸并优化中心位置,以保留从焦点细节到周边环境的层级上下文。大量实验表明,Visual Funnel显著优于单一裁剪和无序多裁剪基线。结果进一步验证,简单增加无序裁剪数量仅带来有限提升甚至负面影响,证明层级结构组合是破解上下文盲区的关键。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) demonstrate impressive reasoning capabilities, but often fail to perceive fine-grained visual details, limiting their applicability in precision-demanding tasks. While methods that crop salient regions of an image offer a partial solution, we identify a critical limitation they introduce: "Contextual Blindness". This failure occurs due to structural disconnect between high-fidelity details (from the crop) and the broader global context (from the original image), even when all necessary visual information is present. We argue that this limitation stems not from a lack of information 'Quantity', but from a lack of 'Structural Diversity' in the model's input. To resolve this, we propose Visual Funnel, a training-free, two-step approach. Visual Funnel first performs Contextual Anchoring to identify the region of interest in a single forward pass. It then constructs an Entropy-Scaled Portfolio that preserves the hierarchical context - ranging from focal detail to broader surroundings - by dynamically determining crop sizes based on attention entropy and refining crop centers. Through extensive experiments, we demonstrate that Visual Funnel significantly outperforms naive single-crop and unstructured multi-crop baselines. Our results further validate that simply adding more unstructured crops provides limited or even detrimental benefits, confirming that the hierarchical structure of our portfolio is key to resolving Contextual Blindness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。