arXiv:2602.03891eess.AScs.AI2026-02

用双路径音频编码器提升视频精彩片段检测效果

Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection

  • 设计双路径音频编码器,分别捕捉语义内容与动态声学特征
  • 在MrHiSum数据集上达到新最优性能,显著提升检测准确率
  • 适合关注多模态视频理解与音频特征建模的研究者

视听视频精彩片段检测旨在结合视觉与听觉线索自动识别视频中最引人注目的时刻。然而,现有模型常低估音频模态的作用,过度依赖高层语义特征,未能充分利用声音的丰富动态特性。为此,本文提出一种新框架——双路径音频编码器用于视频精彩片段检测(DAViHD)。该编码器包含语义路径和动态路径:语义路径通过识别语音、音乐或特定声响事件提取高层信息;动态路径采用随时间演进的频率自适应机制,联合建模声学动态,从而通过显著频带和快速能量变化捕捉瞬时声学事件。将该音频编码器集成至完整的视听框架后,在大规模MrHiSum基准上取得新的最先进性能。结果表明,复杂的双面音频表征是推动该领域发展的关键。

原文摘要 · Abstract (English)

Audio-visual video highlight detection aims to automatically identify the most salient moments in videos by leveraging both visual and auditory cues. However, existing models often underutilize the audio modality, focusing on high-level semantic features while failing to fully leverage the rich, dynamic characteristics of sound. To address this limitation, we propose a novel framework, Dual-Pathway Audio Encoders for Video Highlight Detection (DAViHD). The dual-pathway audio encoder is composed of a semantic pathway for content understanding and a dynamic pathway that captures spectro-temporal dynamics. The semantic pathway extracts high-level information by identifying the content within the audio, such as speech, music, or specific sound events. The dynamic pathway employs a frequency-adaptive mechanism as time evolves to jointly model these dynamics, enabling it to identify transient acoustic events via salient spectral bands and rapid energy changes. We integrate the novel audio encoder into a full audio-visual framework and achieve new state-of-the-art performance on the large-scale MrHiSum benchmark. Our results demonstrate that a sophisticated, dual-faceted audio representation is key to advancing the field of highlight detection.

视频检测音频编码多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。