arXiv:2410.10589cs.CV2024-10NeurIPS被引 10

让视觉语言模型在视频任务中兼顾通用与专用能力

MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer

  • 用动态专家混合机制学习多种任务视角
  • 在Kinetics-400等数据集上实现零样本与闭集性能平衡
  • 适合需要跨任务迁移的视频理解研究者

将大规模视觉语言基础模型的知识迁移到视频识别任务已证明有效。为弥合领域差异,需添加参数化模块以捕捉时序信息。然而,随着专用参数增加,零样本泛化能力下降,现有方法在零样本与闭集性能间存在权衡。本文提出MoTE框架,实现通用性与专用性的统一平衡。通过调优时序专家混合,学习多种任务视角及不同程度的数据拟合。为最大限度保留各专家知识,提出权重合并正则化(Weight Merging Regularization),在权重空间约束专家融合过程;同时引入时序特征调制,规范测试阶段时序特征贡献。在Kinetics-400、Kinetics-600、UCF-101和HMDB-51等多个数据集上,实现零样本与闭集任务的优良平衡,达到或超越当前最优性能。代码已公开于https://github.com/ZMHH-H/MoTE。

原文摘要 · Abstract (English)

Transferring visual-language knowledge from large-scale foundation models for video recognition has proved to be effective. To bridge the domain gap, additional parametric modules are added to capture the temporal information. However, zero-shot generalization diminishes with the increase in the number of specialized parameters, making existing works a trade-off between zero-shot and close-set performance. In this paper, we present MoTE, a novel framework that enables generalization and specialization to be balanced in one unified model. Our approach tunes a mixture of temporal experts to learn multiple task views with various degrees of data fitting. To maximally preserve the knowledge of each expert, we propose \emph{Weight Merging Regularization}, which regularizes the merging process of experts in weight space. Additionally with temporal feature modulation to regularize the contribution of temporal feature during test. We achieve a sound balance between zero-shot and close-set video recognition tasks and obtain state-of-the-art or competitive results on various datasets, including Kinetics-400 \& 600, UCF, and HMDB. Code is available at \url{https://github.com/ZMHH-H/MoTE}.

视频理解知识迁移多专家模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。