arXiv:2412.19563cs.CV2024-12被引 1

用强化学习联合优化音频视频事件分割与标签去噪,提升弱监督下的识别精度。

Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing

  • 基于强化学习构建联合训练框架,让去噪与解析同步优化。
  • 在多个数据集上显著超越传统去噪方法,提升事件边界识别准确率。
  • 适用于各类弱监督音频视频模型,可作为通用增强模块使用。

音频-视觉视频解析(AVVP)旨在识别音频和视觉事件标签并精确划分其时间边界,但因模态中仅提供整体视频标签,任务极具挑战性。现有标签去噪模型通常将去噪作为独立预处理步骤,导致与下游解析任务脱节。为此,我们提出一种基于强化学习的联合标签去噪方法(RLLD),通过联合优化策略实现标签去噪与视频解析模型的同步训练。引入新颖的AVVP验证与软间奖励反馈机制,直接引导标签去噪策略的学习。在多个AVVP任务上的大量实验表明,该方法性能优于现有去噪技术。进一步将本方法集成至其他AVVP模型时,仍能持续提升解析效果。

原文摘要 · Abstract (English)

Audio-visual video parsing (AVVP) aims to recognize audio and visual event labels with precise temporal boundaries, which is quite challenging since audio or visual modality might include only one event label with only the overall video labels available. Existing label denoising models often treat the denoising process as a separate preprocessing step, leading to a disconnect between label denoising and AVVP tasks. To bridge this gap, we present a novel joint reinforcement learning-based label denoising approach (RLLD). This approach enables simultaneous training of both label denoising and video parsing models through a joint optimization strategy. We introduce a novel AVVP-validation and soft inter-reward feedback mechanism that directly guides the learning of label denoising policy. Extensive experiments on AVVP tasks demonstrate the superior performance of our proposed method compared to label denoising techniques. Furthermore, by incorporating our label denoising method into other AVVP models, we find that it can further enhance parsing results.

弱监督视频解析强化学习标签去噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。