arXiv:2411.14401cs.CVcs.LG2024-11ICCV被引 20

不训练也能懂视频,动态合并关键帧提升理解力

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

  • 动态聚类关键帧并选择性压缩令牌,平衡效率与语义
  • 零样本视频理解性能超越现有方法,达成新基准
  • 适合追求高效、无需微调的视频理解场景

多模态大模型为视频理解带来新可能,但零样本任务仍面临高保真度挑战。传统方法依赖微调捕捉时空细节,成本高昂;而无训练方法又常因丢失上下文特征而表现不稳定。为此,我们提出DYTO——一种面向零样本视频理解的动态令牌合并框架,通过分层帧选择与二分图令牌合并策略,自适应地优化令牌效率,同时保留关键场景信息。在多个基准测试中,DYTO显著优于微调与无训练方法,创下零样本视频理解新纪录。

原文摘要 · Abstract (English)

Recent advancements in multimodal large language models (MLLMs) have opened new avenues for video understanding. However, achieving high fidelity in zero-shot video tasks remains challenging. Traditional video processing methods rely heavily on fine-tuning to capture nuanced spatial-temporal details, which incurs significant data and computation costs. In contrast, training-free approaches, though efficient, often lack robustness in preserving context-rich features across complex video content. To this end, we propose DYTO, a novel dynamic token merging framework for zero-shot video understanding that adaptively optimizes token efficiency while preserving crucial scene details. DYTO integrates a hierarchical frame selection and a bipartite token merging strategy to dynamically cluster key frames and selectively compress token sequences, striking a balance between computational efficiency with semantic richness. Extensive experiments across multiple benchmarks demonstrate the effectiveness of DYTO, achieving superior performance compared to both fine-tuned and training-free methods and setting a new state-of-the-art for zero-shot video understanding.

视频理解零样本动态合并多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。