arXiv:2607.12557cs.CV2026-07中稿 · PRCV 2026

用高斯混合模型精准选帧,节省一半算力还更准

Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding

论文配图:Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding
图 1 · 摘自论文原文
  • 用高斯混合模型捕捉事件结构,智能区分主次帧
  • 每事件保留1张高清主帧+低清辅帧,仅用一半令牌预算
  • 不需训练、适配多模型,长视频理解更高效

大视觉语言模型在长视频理解中面临计算成本过高和信息丢失问题,现有关键帧选择方法常将帧视为原子单元并均等分配视觉预算,忽视高层语义结构且冗余严重。为此,我们提出GMM-EVA(基于高斯混合模型的事件感知视觉分配),通过高斯混合模型从离散帧观测中建模事件级结构。采用差异化分配策略:每事件保留1张高分辨率主关键帧以保持细节清晰度,同时使用低分辨率次关键帧维持时间上下文并优化令牌预算。GMM-EVA为无训练、即插即用框架,在多种相关性度量和下游LVLM上具有强泛化能力。大量实验表明,该方法显著优于均匀采样;值得注意的是,其性能接近基线选择方法,却仅需约一半视觉令牌预算,展现出卓越的效率与有效性。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) face significant challenges in long video understanding due to the excessive computational cost and information loss associated with uniform sampling. Existing keyframe selection methods often treat video frames as atomic entities and allocate visual budgets equally, thereby overlooking high-level semantic structures and introducing substantial redundancy. To address these limitations, we propose GMM-EVA (Gaussian Mixture Modeling for Event-Aware Visual Allocation), which leverages Gaussian Mixture Models to model event-level structure from discrete frame-wise observations. A differentiated allocation strategy is then applied to preserve one primary high-resolution keyframe per event for high-fidelity detail, while utilizing lower-resolution secondary keyframes to maintain temporal context and optimize token budgets. GMM-EVA is a training-free, plug-and-play framework that generalizes robustly across various relevance measures and downstream LVLMs. Extensive experiments on multiple long video benchmarks demonstrate that our method significantly outperforms uniform sampling. Notably, GMM-EVA achieves comparable performance to baseline selection methods while utilizing only approximately half of the visual token budget, highlighting its superior efficiency and effectiveness.

长视频理解关键帧选择高斯混合模型视觉预算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。