arXiv:2608.20473cs.CV2026-08

用最优传输方法压缩视频令牌,保留关键视觉信息。

Aggregating Visual Information with Optimal Transport for VideoLM Token Compression

论文配图:Aggregating Visual Information with Optimal Transport for VideoLM Token Compression
图 1 · 摘自论文原文
  • 将视频令牌压缩建模为从密集观测到紧凑目标的最优传输问题。
  • 在多种压缩率下表现优于或匹配未压缩基线,高压缩时仍保持性能。
  • 支持任务条件和空间粒度自适应,适合视频理解场景。

视频语言模型将视频处理为稠密的视觉令牌序列,存在显著的表示冗余。压缩这些序列对于减轻语言模型解码时的视觉令牌负担至关重要。核心挑战在于如何在压缩过程中保留分散在各帧中的视觉信息。为此,我们提出基于最优传输的视觉信息聚合方法(AVIOT),将视频令牌压缩建模为将帧观测的密集经验测度运输到一个紧凑的目标测度。由此产生的源到目标耦合为每个目标支撑点定义了源观测的分布,直接决定了压缩后视频表示的构建方式。进一步地,该方法沿任务和空间轴进行适应:问题条件调节源帧与目标支撑点间的传输代价,影响各时间片段分配的支撑点数量,从而引导表示能力聚焦于相关内容;在多尺度空间粒度下,AVIOT计算区域特异的时空传输计划,并自适应融合其生成的表示,使同一紧凑表示中的不同区域可提取不同时刻的信息。在多种压缩比下的评估显示,AVIOT在多个视频理解基准上达到或超越未压缩基线表现,且在更高压缩率下仍保持优异性能。

原文摘要 · Abstract (English)

Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing the visual-token burden on language-model decoding. The central challenge is to preserve visual information dispersed across frames under such compression. To this end, we introduce Aggregating Visual Information with Optimal Transport (AVIOT), which casts video token compression as transporting a dense empirical measure of frame observations onto a compact target measure. The resulting source-to-target coupling induces a distribution over source observations for each target support, directly specifying how the compressed video representation is constructed. We further adapt this construction along task and spatial axes. Question conditioning modulates the transport cost between source frames and target supports, while influencing how many supports are allocated to each temporal segment, thereby directing representation capacity toward question-relevant content. At multiple spatial granularities, AVIOT computes region-specific temporal transport plans and adaptively fuses the representations they yield, allowing different regions within the same compact representation to draw from different moments. Evaluations across varying compression ratios show that AVIOT matches or outperforms the uncompressed baseline on multiple video-understanding benchmarks while retaining strong performance at higher compression ratios.

视频理解最优传输令牌压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。