首个面向竖屏短视频的音视频事件定位数据集,解决手机时代多模态理解新挑战
Audio-visual Event Localization on Portrait Mode Short Videos
- 构建首个竖屏短视音频事件定位数据集,含25,335个带帧级标注的片段
- 现有方法在竖屏视频上性能平均下降18.66%,因画面方向与复杂音频干扰
- 提出适配竖屏视频的预处理和建模方案,为移动端多模态研究提供新思路
音视频事件定位(AVEL)在多模态场景理解中至关重要。现有AVEL数据集多为横屏长视频,音频环境清晰简单,而智能手机普及使短视频成为主流内容形式。这类视频具有竖屏构图和分层音频(如重叠音效、旁白与背景音乐)的特点,带来传统方法未覆盖的新挑战。为此,我们提出AVE-PM,首个专为竖屏短视频设计的AVEL数据集,包含25,335个片段,覆盖86个细粒度类别,且具帧级标注。实证分析显示,先进AVEL方法在跨模式评估中平均性能下降18.66%。进一步分析揭示两大核心挑战:1)竖屏构图引入空间偏置,形成独特领域先验;2)噪声音频组合降低音频模态可靠性。我们探究了最优预处理方案及背景音乐的影响,实验表明针对性预处理与专用模型设计仍可显著提升性能。本工作为移动主导视频时代的AVEL研究提供了基准与实用洞见。数据集与代码将公开。
原文摘要 · Abstract (English)
Audio-visual event localization (AVEL) plays a critical role in multimodal scene understanding. While existing datasets for AVEL predominantly comprise landscape-oriented long videos with clean and simple audio context, short videos have become the primary format of online video content due to the the proliferation of smartphones. Short videos are characterized by portrait-oriented framing and layered audio compositions (e.g., overlapping sound effects, voiceovers, and music), which brings unique challenges unaddressed by conventional methods. To this end, we introduce AVE-PM, the first AVEL dataset specifically designed for portrait mode short videos, comprising 25,335 clips that span 86 fine-grained categories with frame-level annotations. Beyond dataset creation, our empirical analysis shows that state-of-the-art AVEL methods suffer an average 18.66% performance drop during cross-mode evaluation. Further analysis reveals two key challenges of different video formats: 1) spatial bias from portrait-oriented framing introduces distinct domain priors, and 2) noisy audio composition compromise the reliability of audio modality. To address these issues, we investigate optimal preprocessing recipes and the impact of background music for AVEL on portrait mode videos. Experiments show that these methods can still benefit from tailored preprocessing and specialized model design, thus achieving improved performance. This work provides both a foundational benchmark and actionable insights for advancing AVEL research in the era of mobile-centric video content. Dataset and code will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。