用频率残差压缩视频令牌,兼顾细节与时序覆盖。
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs

- 分离空间与时间信息:保留稀疏高保真空间锚点,用频域残差表示密集时序变化。
- 在视觉潜空间中对帧间残差做1D-DCT,低频集中显著压缩令牌长度。
- 适合需要长视频理解且资源受限的多模态大模型应用。
视频多模态大模型面临空间保真度与时间覆盖范围之间的持续矛盾:保留精细视觉细节需要大量空间令牌,而捕捉短暂事件则需密集的时间采样。本文提出 extbf{Fre-Res},一种预算自适应的双轨视频令牌压缩框架,将这两类证据分离处理。Fre-Res 保留稀疏的高保真空间锚点,并通过紧凑的残差-频率令牌表示密集的时序演化。具体地,它在视觉潜空间对帧间残差轨迹进行一维离散余弦变换(1D-DCT),观察到强低频集中现象。为使频域动态与原始视觉嵌入对齐,Fre-Res 引入空间引导吸收器,将时序残差信息注入对应的空间锚点令牌。在细粒度短视频与长视频推理基准上,Fre-Res 实现了良好的精度-效率权衡,达到或接近全令牌性能,同时大幅减少视觉令牌长度。大量消融实验进一步表明,时序-频率残差保留因果过渡线索,而空间锚点对精细物体与布局推理仍至关重要。
原文摘要 · Abstract (English)
Video MLLMs face a persistent tension between spatial fidelity and temporal coverage: preserving fine-grained visual details requires many spatial tokens, while capturing short-lived events requires dense temporal sampling. We propose \textbf{Fre-Res}, a budget-adaptive dual-track video-token compression framework that separates these two forms of evidence. Fre-Res preserves sparse high-fidelity spatial anchors and represents dense temporal evolution through compact residual-frequency tokens. Specifically, it applies temporal 1D-DCT to inter-frame residual trajectories in vision-latent space, where we observe strong low-frequency concentration. To align frequency-domain dynamics with native visual embeddings, Fre-Res introduces a Spatial-Guided Absorber that injects temporal residual information into spatially corresponding anchor tokens. Across fine-grained short-video and long-video reasoning benchmarks, Fre-Res achieves a favorable accuracy--efficiency trade-off, matching or approaching full-token performance while substantially reducing visual-token length. Extensive ablations further show that temporal-frequency residuals preserve causal transition cues, while spatial anchors remain essential for fine-grained object and layout reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。