用多模态大模型直接提升少样本动作识别,效果更好且参数更少。
Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs

- 利用大模型生成时空语义丰富的多模态特征,分离优化视觉与文本表示。
- 在多个数据集上达到新最佳,仅需极少可训练参数。
- 适合做少样本学习、多模态理解的研究者和应用开发者。
多模态大语言模型(MLLMs)推动了少样本动作识别(FSAR)的发展。然而,现有方法主要依赖生成描述文本,形成低效的特征→文本→特征流程,并仅在视觉空间内进行度量学习。本文提出首个端到端方法FSAR-LLaVA,将Video-LLaVA等MLLM作为多模态知识库,直接增强FSAR。首先,在特征层面,利用MLLM的多模态解码器提取时空与语义融合的表示,通过提出的多模态特征增强模块解耦并增强为独立的视觉与文本特征,充分挖掘其语义信息。其次,借助MLLM的灵活性设计适应多种场景的输入提示,利用其对齐输出驱动我们设计的复合任务导向原型构建,有效弥合元训练与元测试分布差异。最后,为实现多模态特征联合指导度量学习,提出无需训练的多模态原型匹配度量,自适应选择关键线索,高效利用MLLM生成的解耦特征表示。大量实验表明,该方法在各类任务中表现优异,且可训练参数极少。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have propelled the field of few-shot action recognition (FSAR). However, preliminary explorations in this area primarily focus on generating captions to form a suboptimal feature->caption->feature pipeline and adopt metric learning solely within the visual space. In this paper, we propose FSAR-LLaVA, the first end-to-end method to leverage MLLMs (such as Video-LLaVA) as a multimodal knowledge base for directly enhancing FSAR. First, at the feature level, we leverage the MLLM's multimodal decoder to extract spatiotemporally and semantically enriched representations, which are then decoupled and enhanced by our Multimodal Feature-Enhanced Module into distinct visual and textual features that fully exploit their semantic knowledge for FSAR. Next, we leverage the versatility of MLLMs to craft input prompts that flexibly adapt to diverse scenarios, and use their aligned outputs to drive our designed Composite Task-Oriented Prototype Construction, effectively bridging the distribution gap between meta-train and meta-test sets. Finally, to enable multimodal features to guide metric learning jointly, we introduce a training-free Multimodal Prototype Matching Metric that adaptively selects the most decisive cues and efficiently leverages the decoupled feature representations produced by MLLMs. Extensive experiments demonstrate superior performance across various tasks with minimal trainable parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。