通过分层语义约束图提升跨模态事件定位精度
Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization

- 构建异构分层图,融合音视频段与视频级节点
- 在开放词汇场景下达到新最优,显著提升定位准确率
- 适合多模态理解、视频分析方向研究者参考
开放词汇音视频事件定位(OV-AVEL)需联合建模音视频线索以识别并定位训练中未见的事件。现有方法多在欧氏空间学习联合表示,仍面临两大挑战:一是未见类别缺乏监督信号,难以维持多时间尺度的音视频一致性;二是段级与视频级语义间缺乏层次约束,导致跨层级语义不一致。为此,本文提出分层语义约束异构图(HSCHG)框架。首先在欧氏空间构建包含音频/视觉段节点及其对应视频级节点的异构分层图,使用多向时间边捕捉模态内完整时序信息;同时引入双阈值过滤门控融合策略,仅在对齐置信度高时融合跨模态信息。进一步设计段级与视频级表示间的双向语义约束,实现多层级语义一致性。基于此,将多层级音视频表示与文本原型统一映射至双曲空间,并采用分层蕴含正则化损失刻画视频与段之间的层次关系。大量实验表明,本方法在OV-AVEL基准上优于现有方法,消融实验验证了其有效性。
原文摘要 · Abstract (English)
Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including categories unseen during training. Existing methods primarily learn joint audio-visual representations in Euclidean space, but still face two significant challenges. First, the lack of supervision signals for unseen categories makes it difficult to maintain audio-visual consistency across multiple temporal scales. Second, the lack of hierarchical constraints between segment- and video-level semantics prevents the model from establishing semantic consistency across different levels. To address these challenges, we propose a hierarchical semantic constrained heterogeneous graph (HSCHG) for audio-visual event localization framework. We first construct a heterogeneous hierarchical graph in Euclidean space, which includes audio and visual segment nodes and their corresponding video-level nodes. We use multi-directional temporal edges to capture complete temporal information within each modality. Simultaneously, we employ a dual-threshold filtering gated fusion strategy, introducing cross-modal information only when the alignment confidence is high. Furthermore, we introduce bidirectional semantic constraints between segment- and video-level representations to achieve semantic consistency across different levels. Based on this, we map the multi-level audio-visual representations and text prototypes uniformly into hyperbolic space. We use a hierarchical entailment regularization loss to characterize the hierarchical relationships between videos and segments. Extensive experimental results show that our method outperforms existing methods on the OV-AVEL benchmark. Ablation studies further validate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。