用设施选址法高效压缩长视频视觉令牌,提升处理速度。
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
- 基于设施选址函数选择最具代表性且多样化的视觉令牌子集。
- 在有限令牌数预算下,压缩率高且性能接近最优。
- 无需训练、适配多种视频大模型,适合实际部署。
近期长视频理解研究利用大型多模态模型(LMMs)的先进视觉-语言推理能力,推动了专用于处理长视频序列的视频-LMM发展。然而,这些模型的可扩展性严重受限于长视频序列生成的海量视觉令牌。为此,我们提出FLoC,一种基于设施选址函数的高效视觉令牌压缩框架,该方法能快速在预设令牌数量预算内,选取紧凑但高度代表性与多样性的视觉令牌子集。通过集成懒惰贪心算法,我们的方法显著提升效率,大幅减少视觉令牌数量,同时保证近似最优性能。值得注意的是,该方法无需训练、模型无关、查询无关,可无缝集成至多种视频-LLMs及现有工作流中。在Video-MME、MLVU、LongVideoBench和EgoSchema等大规模基准上的广泛评估表明,本框架持续优于近期压缩技术,凸显其在长视频理解中的有效性、鲁棒性及处理效率。
原文摘要 · Abstract (English)
Recent studies in long video understanding have harnessed the advanced visual-language reasoning capabilities of Large Multimodal Models (LMMs), driving the evolution of video-LMMs specialized for processing extended video sequences. However, the scalability of these models is severely limited by the overwhelming volume of visual tokens generated from extended video sequences. To address this challenge, we propose FLoC, an efficient visual token compression framework based on the facility location function, a principled approach that swiftly selects a compact yet highly representative and diverse subset of visual tokens within a predefined budget on the number of visual tokens. By integrating the lazy greedy algorithm, our method achieves remarkable efficiency gains by swiftly selecting a compact subset of tokens, drastically reducing the number of visual tokens while guaranteeing near-optimal performance. Notably, our approach is training-free, model-agnostic, and query-agnostic, providing a versatile solution that seamlessly integrates with diverse video-LLMs and existing workflows. Extensive evaluations on large-scale benchmarks, such as Video-MME, MLVU, LongVideoBench, and EgoSchema, show that our framework consistently surpasses recent compression techniques, highlighting its effectiveness and robustness in addressing the challenges of long video understanding as well as its processing efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。