arXiv:2502.00426cs.CV2025-02IJCAI被引 4

通过动态调整支持集提升零样本视频分类性能

TEST-V: TEst-time Support-set Tuning for Zero-shot Video Classification

  • 用多提示扩展支持集,再通过可学习权重动态筛选关键特征
  • 在四个基准上达到最优效果,支持集调整过程可解释
  • 适合关注零样本视频理解与模型可解释性的研究者

近期通过微调类别嵌入或用生成的视觉样本替换类别名来适应视觉语言模型进行零样本图像分类取得了良好效果。然而,提示调优无法消除模态间的语义鸿沟,而支持集不可调优。为此,我们结合两者优势,提出一种新框架TEST-V:首先利用LLM生成的多个提示对支持集进行扩张(多提示支持集扩张,MSD),丰富支持集多样性;然后通过可学习权重对支持集进行侵蚀(时序感知支持集侵蚀,TSE),依据时序预测一致性自监督地挖掘每个类别的关键支持线索。TEST-V在四个基准上均取得当前最优性能,且支持集的扩张与侵蚀过程具有良好的可解释性。

原文摘要 · Abstract (English)

Recently, adapting Vision Language Models (VLMs) to zero-shot visual classification by tuning class embedding with a few prompts (Test-time Prompt Tuning, TPT) or replacing class names with generated visual samples (support-set) has shown promising results. However, TPT cannot avoid the semantic gap between modalities while the support-set cannot be tuned. To this end, we draw on each other's strengths and propose a novel framework namely TEst-time Support-set Tuning for zero-shot Video Classification (TEST-V). It first dilates the support-set with multiple prompts (Multi-prompting Support-set Dilation, MSD) and then erodes the support-set via learnable weights to mine key cues dynamically (Temporal-aware Support-set Erosion, TSE). Specifically, i) MSD expands the support samples for each class based on multiple prompts enquired from LLMs to enrich the diversity of the support-set. ii) TSE tunes the support-set with factorized learnable weights according to the temporal prediction consistency in a self-supervised manner to dig pivotal supporting cues for each class. $\textbf{TEST-V}$ achieves state-of-the-art results across four benchmarks and has good interpretability for the support-set dilation and erosion.

零样本分类视频理解支持集调优可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。