用视觉与注意力熵动态聚焦关键区域,加速视觉语言动作模型推理。
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
- 通过图像熵和注意力熵识别重要视觉与语义区域。
- 推理速度提升显著,参数量减少且性能优于现有方法。
- 无需训练即可部署,适合实时性要求高的应用。
视觉-语言-动作(VLA)模型融合视觉感知、语言理解与动作决策,实现跨模态语义对齐,具有广泛应用潜力。然而,高维视觉特征、复杂语言输入与连续动作序列的联合处理带来巨大计算开销,导致推理效率低下,阻碍实时部署与可靠性。为此,本文利用图像熵量化每个视觉标记的灰度分布特性,引入注意力熵捕捉任务相关文本的注意力分数分布。视觉熵可识别纹理丰富或结构信息丰富的区域,注意力熵则定位语义相关的标记。结合时间步信息,该方法构建动态切换策略,引导模型从全局视觉特征转向注意力引导的局部关键区域。由此产生的VLA-InfoEntropy方法融合空间、语义与时间线索,在降低冗余的同时保留关键内容。大量实验表明,该方法有效减少推理参数、加速推理速度,并优于现有方法。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models integrate visual perception, language understanding, and action decision-making for cross-modal semantic alignment, exhibiting broad application potential. However, the joint processing of high-dimensional visual features, complex linguistic inputs, and continuous action sequences incurs significant computational overhead and low inference efficiency, thereby hindering real-time deployment and reliability. To address this issue, we use image entropy to quantify the grayscale distribution characteristics of each visual token and introduce attention entropy to capture the distribution of attention scores over task-related text. Visual entropy identifies texture-rich or structurally informative regions, while attention entropy pinpoints semantically relevant tokens. Combined with timestep information, these metrics enable a dynamic transition strategy that shifts the model's focus from global visual features to attention-guided local informative regions. Thus, the resulting VLA-InfoEntropy method integrates spatial, semantic, and temporal cues to reduce redundancy while preserving critical content. Extensive experiments show that our method reduces inference parameters, accelerates inference speed, and outperforms existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。