arXiv:2506.13322cs.CVcs.AI2025-06IJCAI被引 3

主动选择可靠模态,提升少样本动作识别准确率

Active Multimodal Distillation for Few-shot Action Recognition

  • 基于上下文线索动态选择每样本的可靠模态
  • 在多个基准上显著超越现有方法
  • 适合少样本场景下多模态融合任务

由于进展迅速且应用前景广阔,少样本动作识别受到广泛关注。然而,当前方法主要依赖有限的单模态数据,未能充分利用多模态信息潜力。本文提出一种新框架,通过任务特定上下文线索主动识别每个样本的可靠模态,显著提升识别性能。框架集成主动样本推理(ASI)模块,利用主动推理基于后验分布预测可靠模态并进行组织。不同于强化学习,主动推理以证据偏好替代奖励,实现更稳定预测。此外,引入主动互蒸馏模块,通过将可靠模态知识迁移至较不可靠模态来增强表征学习。元测试阶段采用自适应多模态推理,为可靠模态分配更高权重。在多个基准上的大量实验表明,该方法显著优于现有方法。

原文摘要 · Abstract (English)

Owing to its rapid progress and broad application prospects, few-shot action recognition has attracted considerable interest. However, current methods are predominantly based on limited single-modal data, which does not fully exploit the potential of multimodal information. This paper presents a novel framework that actively identifies reliable modalities for each sample using task-specific contextual cues, thus significantly improving recognition performance. Our framework integrates an Active Sample Inference (ASI) module, which utilizes active inference to predict reliable modalities based on posterior distributions and subsequently organizes them accordingly. Unlike reinforcement learning, active inference replaces rewards with evidence-based preferences, making more stable predictions. Additionally, we introduce an active mutual distillation module that enhances the representation learning of less reliable modalities by transferring knowledge from more reliable ones. Adaptive multimodal inference is employed during the meta-test to assign higher weights to reliable modalities. Extensive experiments across multiple benchmarks demonstrate that our method significantly outperforms existing approaches.

少样本学习多模态动作识别主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。