arXiv:2509.14893cs.SDeess.AS2025-09

通过异质时间图对比学习,提升多模态声音事件分类精度

Temporally Heterogeneous Graph Contrastive Learning for Multimodal Acoustic event Classification

  • 构建事件级时序图,区分模态内平滑与模态间衰减
  • 在AudioSet上达到当前最佳性能,显著降低噪声干扰
  • 适合音频视觉融合、事件识别等场景研究者

多模态声音事件分类在音视频系统中至关重要。尽管融合音视频信号可提升识别效果,但时序对齐困难且跨模态噪声影响大。现有方法常独立处理音视频流,后期通过对比或互信息融合特征。近期多模态图学习虽有进展,但多数未能区分模态内与模态间的时间依赖。为此,我们提出时序异质图对比学习(THGCL)。该框架为每个事件构建时序图,以音频和视频片段为节点,时序连接为边;引入高斯过程建模模态内平滑性,霍克斯过程建模模态间衰减,并结合对比学习捕捉细粒度关系。在AudioSet数据集上的实验表明,THGCL取得当前最优性能。

原文摘要 · Abstract (English)

Multimodal acoustic event classification plays a key role in audio-visual systems. Although combining audio and visual signals improves recognition, it is still difficult to align them over time and to reduce the effect of noise across modalities. Existing methods often treat audio and visual streams separately, fusing features later with contrastive or mutual information objectives. Recent advances explore multimodal graph learning, but most fail to distinguish between intra- and inter-modal temporal dependencies. To address this, we propose Temporally Heterogeneous Graph-based Contrastive Learning (THGCL). Our framework constructs a temporal graph for each event, where audio and video segments form nodes and their temporal links form edges. We introduce Gaussian processes for intra-modal smoothness, Hawkes processes for inter-modal decay, and contrastive learning to capture fine-grained relationships. Experiments on AudioSet show that THGCL achieves state-of-the-art performance.

多模态声音事件图学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。