动态对齐+渐进细节增强,提升视觉语言模型精度与效率
DAPE: Dynamic Non-uniform Alignment and Progressive Detail Enhancement Techniques for Improving the Performance of Efficient Visual Language Models

- 动态匹配机制按信息密度分配图像块数量与大小
- 渐进引入高分辨率特征,准确率显著提升且计算开销降低
- 适合追求高效高精度多模态模型的开发者与研究者
近年来,预训练视觉-语言模型展现出巨大潜力,成为众多下游任务的基础框架。然而,文本与图像间的信息密度分布不均,现有方法常忽视文本标签与图像块在信息密度和语义范围上的动态差异,统一对齐策略导致跨模态交互粗粒度且丢失细粒度语义。此外,更精细对齐通常伴随巨大计算开销,限制实际部署。为此,本文提出一种动态交叉模态对齐与连续细节增强框架。首先设计可学习的动态匹配机制,根据文本标签的信息密度动态分配不同数量与尺寸的图像标签,实现更精准的关注交互。其次构建连续细节引入模块,逐步将高分辨率视觉特征融入对齐过程。在多个基准测试上的大量实验表明,该方法显著提升了各类下游任务的准确率,同时降低计算开销。
原文摘要 · Abstract (English)
In recent years, pre-trained visual-linguistic models have demonstrated tremendous potential, becoming a crucial foundational framework for numerous downstream tasks. However, the information density between text and images is not uniformly distributed. Existing methods often overlook the inherent and dynamic differences in information density and semantic scope between text tags and image blocks. These common uniform alignment strategies result in coarse-grained cross-modal interactions and loss of fine semantic details. Moreover, pursuing finer alignment typically requires substantial computational overhead, limiting practical model deployment. To address this challenge, this paper proposes a novel framework for dynamic cross-modal alignment with continuous detail introduction. First, we design a dynamically adaptive cross-modal matching mechanism that uses a learnable matching function to dynamically assign varying numbers and sizes of image tags to text tags of the same size but different information density, enabling more precise attention interaction. Second, we develop a continuous detail introduction module to progressively incorporate high-resolution visual feature enhancement into the alignment process. Extensive experiments across multiple benchmarks demonstrate significant improvements in the accuracy of various downstream tasks while reducing computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。