arXiv:2509.04086cs.CVcs.MM2025-09被引 1

通过文本增强与类别感知图结构,提升弱监督音视频事件分割精度

TEn-CATG:Text-Enriched Audio-Visual Video Parsing with Multi-Scale Category-Aware Temporal Graph

  • 用双向文本融合模块对齐跨模态语义,减少伪标签噪声
  • 引入多尺度类别感知时序图,实现事件相关的时序推理
  • 在LLP和UnAV-100数据集上均达最优,适合弱监督视频理解任务

音视频视频解析(AVVP)旨在弱监督条件下检测视频中的事件类别及其时间边界。现有方法主要聚焦于使用基于注意力的架构改进时序建模,或生成更丰富的伪标签以弥补帧级标注缺失。然而,注意力模型常因过拟合噪声伪标签导致误差累积,而伪标签生成方法则在帧间均匀分配注意力,削弱了时序定位精度。为此,本文提出TEn-CATG框架,结合语义校准与类别感知时序推理。具体地,设计双向文本融合(BiT)模块,利用音视频特征作为语义锚点来优化文本嵌入,区别于传统文本到特征对齐方式,从而降低噪声并增强跨模态一致性。同时引入类别感知时序图(CATG)模块,通过选择多尺度时序邻居并学习类别特异性时序衰减因子,实现有效的事件依赖时序推理。大量实验表明,TEn-CATG在基准数据集LLP和UnAV-100上多个评估指标上达到最新水平,凸显其在弱监督AVVP任务中捕捉复杂时序与语义依赖关系的鲁棒性与优越性。

原文摘要 · Abstract (English)

Audio-visual video parsing (AVVP) aims to detect event categories and their temporal boundaries in videos, typically under weak supervision. Existing methods mainly focus on (i) improving temporal modeling using attention-based architectures or (ii) generating richer pseudo-labels to address the absence of frame-level annotations. However, attention-based models often overfit noisy pseudo-labels, leading to cumulative training errors, while pseudo-label generation approaches distribute attention uniformly across frames, weakening temporal localization accuracy. To address these challenges, we propose TEn-CATG, a text-enriched AVVP framework that combines semantic calibration with category-aware temporal reasoning. More specifically, we design a bi-directional text fusion (BiT) module by leveraging audio-visual features as semantic anchors to refine text embeddings, which departs from conventional text-to-feature alignment, thereby mitigating noise and enhancing cross-modal consistency. Furthermore, we introduce the category-aware temporal graph (CATG) module to model temporal relationships by selecting multi-scale temporal neighbors and learning category-specific temporal decay factors, enabling effective event-dependent temporal reasoning. Extensive experiments demonstrate that TEn-CATG achieves state-of-the-art results across multiple evaluation metrics on benchmark datasets LLP and UnAV-100, highlighting its robustness and superior ability to capture complex temporal and semantic dependencies in weakly supervised AVVP tasks.

音视频解析弱监督时序建模跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。