arXiv:2605.25952cs.CVcs.AI2026-05

通过融合多视角视觉表征提升效率与性能,解决压缩导致信息丢失问题。

VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding

论文配图:VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding
图 1 · 摘自论文原文
  • 先融合多视角视觉信息再自适应压缩,提升特征密度。
  • 在少令牌情况下仍保持高精度,复杂任务表现更优。
  • 适合需要高效高精度多模态理解的场景。

尽管近期高效方法在加速多模态理解方面取得显著进展,但仍面临明显性能下降问题。其对单一视觉线索的高压缩率和依赖粗粒度注意力对齐的启发式剪枝策略,限制了视觉标记的信息容量与密度。为此,我们提出VEN-VL,一种遵循‘先丰富后压缩’原则的视觉集成混合专家(MoE)框架,用于高效且有效的感知。具体而言,我们首先通过统一不同视角的视觉表征来增强信息容量,然后通过专用视觉专家中的自适应路由器逐步压缩,以提高信息密度。此外,我们通过显式视觉监督引入原始结构的重建能力,有助于关键信息的保留。实验结果表明,我们在使用少量信息浓缩标记的情况下,在复杂视觉任务中展现出优越性,有效弥合了性能与效率之间的差距。

原文摘要 · Abstract (English)

Despite the remarkable progress achieved by recent efficient methods in accelerating multimodal understanding, they still suffer from noticeable performance degradation. Their emphasis on the high compression ratio of a single visual clue and reliance on the heuristic pruning strategy with coarse attention alignment incurs a bottleneck on the information capacity and density of visual tokens. Addressing this limitation, we propose VEN-VL, a visual ensemble MoE framework for effective and efficient perception following the enrich then compact principle. Specifically, we first enrich the information capacity by unifying the visual representations of different perspectives, and then progressively compact it with adaptive routers in specialized visual experts to enhance the information density. Furthermore, we incorporate the reconstruction ability of vanilla structure via explicit visual supervision, facilitating crucial information preservation. Experimental results demonstrate our superiority in complex visual tasks with few information-condensed tokens, which effectively bridges the gap between performance and efficiency.

多模态MoE视觉压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。