arXiv:2507.09323cs.CV2025-07

让音视频模型更好区分易混淆动作,提升识别准确率。

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition

  • 动态调整混淆损失,强化相似类别的区分能力。
  • 在VGGSound上达到65.5%的顶级准确率。
  • 适合需要精细动作识别的多模态应用。

人类理解事件时不会孤立看待,而是通过类内泛化与类间对比进行认知。现有音视频预训练方法仅关注整体模态对齐,未考虑通过认知引导和对比学习增强易混淆类别的区分能力。本文提出动态跨类别混淆感知编码器(DICCAE),在细粒度类别层级实现音视频表征对齐。DICCAE通过动态调节基于类别间混淆程度的损失,提升模型对相似活动的辨别能力。为进一步拓展应用,引入融合音频、视频及其融合表示的新型训练框架。针对人体动作识别中音视频数据稀缺问题,提出聚类引导的自监督预训练策略。在VGGSound数据集上,DICCAE实现65.5%的top-1准确率,接近当前最优水平。通过大量消融实验验证各模块必要性,证明其表征质量优越。

原文摘要 · Abstract (English)

Humans do not understand individual events in isolation; rather, they generalize concepts within classes and compare them to others. Existing audio-video pre-training paradigms only focus on the alignment of the overall audio-video modalities, without considering the reinforcement of distinguishing easily confused classes through cognitive induction and contrast during training. This paper proposes the Dynamic Inter-Class Confusion-Aware Encoder (DICCAE), an encoder that aligns audio-video representations at a fine-grained, category-level. DICCAE addresses category confusion by dynamically adjusting the confusion loss based on inter-class confusion degrees, thereby enhancing the model's ability to distinguish between similar activities. To further extend the application of DICCAE, we also introduce a novel training framework that incorporates both audio and video modalities, as well as their fusion. To mitigate the scarcity of audio-video data in the human activity recognition task, we propose a cluster-guided audio-video self-supervised pre-training strategy for DICCAE. DICCAE achieves near state-of-the-art performance on the VGGSound dataset, with a top-1 accuracy of 65.5%. We further evaluate its feature representation quality through extensive ablation studies, validating the necessity of each module.

音视频融合动作识别自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。