提出频率增强方法,让视觉语言模型更准识别细粒度图像信息
HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models

- 引入任务感知的频域路径,动态获取多层级频率特征
- 在多个数据集上提升细粒度理解与幻觉抵抗能力,效果优于现有方法
- 适合需要精准视觉推理的应用,如医学图像分析、复杂场景问答
视觉语言模型(VLMs)在依赖细粒度视觉证据的任务中仍不可靠。我们发现一个被忽视的原因:频谱响应僵化。尽管图像和任务间存在显著频率差异,预训练视觉编码器在各层仍保持固定的频谱特征分布,微调时变化极小。因预训练仅接触图像,无法根据当前问题调整频谱提取。为此,我们提出HAFI-VLM,引入任务条件频域路径,同时保留预训练语义表示。层次自适应频率注入(HAFI)通过文本调制的跨注意力,在多层编码器中获取互补的低、中、高频证据。视觉增强层适配器重新校准浅层大模型注意力,有效利用增强后的视觉令牌。在LLaVA-1.5和Qwen2.5-VL上的实验表明,该方法在通用视觉问答、文本丰富理解及幻觉鲁棒性上均有持续提升,优于表示层增强方法和多数分辨率或裁剪类方法,且无需额外高分辨率编码。机制分析显示,HAFI恢复了任务相关的频谱分配,同时保留语义注意力,确立频谱增强为提升VLM感知力的独立有效路径。
原文摘要 · Abstract (English)
Vision-language models (VLMs) remain unreliable when predictions require fine-grained visual evidence. We identify a previously overlooked cause: spectral response rigidity. Despite substantial frequency variation across images and tasks, pretrained vision encoders exhibit persistent, encoder-specific layerwise spectral profiles that change only marginally under downstream fine-tuning. Since pretrained vision encoders only receive images, they cannot adapt spectral extraction to the evidence required by the current query. We therefore propose HAFI-VLM, which introduces a task-conditioned frequency pathway while preserving the pretrained semantic representation. Hierarchical Adaptive Frequency Injection (HAFI) retrieves complementary low-, mid-, and high-frequency evidence at multiple encoder depths using text-modulated, spatially aligned cross-attention. A Visual Enrichment Layer Adapter further recalibrates shallow LLM attention to effectively utilize the enriched visual tokens. Experiments on LLaVA-1.5 and Qwen2.5-VL demonstrate consistent improvements in general VQA, text-rich understanding, and hallucination robustness, outperforming representation-level enhancement methods and most resolution- or cropping-based approaches without additional high-resolution encoding. Mechanistic analyses show that HAFI restores task-dependent spectral allocation while retaining semantic attention, establishing frequency enrichment as a distinct and effective route for improving VLM perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。