统一处理音视频场景理解任务,让模型像人一样综合感知
AV-Unified: A Unified Framework for Audio-visual Scene Understanding
- 将不同任务统一为离散标记序列,共享同一架构
- 多尺度时序模块捕捉事件时间差异,跨模态引导建模空间关联
- 适合需要多任务协同理解音视频的科研与应用
人类感知世界时会自然融合多种音视频任务。但现有研究如事件定位、解析、分割和问答多为独立探索,难以全面理解复杂音视频场景并挖掘任务间关系。为此,我们提出AV-Unified,一个支持多种音视频场景理解任务的统一框架。该框架通过将各任务输入输出转化为离散标记序列,标准化格式并建立共享表示,使单一模型可在异构数据集上联合训练。针对音视频事件的时间粒度差异,设计多尺度时序感知模块以捕捉关键线索;为弥补视觉域中缺乏听觉监督的问题,引入基于跨模态引导的空间感知模块,建模空间上的音视频关联。此外,采用任务特定文本提示增强模型适应性与任务感知能力。在多个基准数据集(如AVE、LLP、MUSIC-AVQA、VGG-SS和AVS)上的大量实验表明,AV-Unified在时间、空间及时空任务上均表现优异。
原文摘要 · Abstract (English)
When humans perceive the world, they naturally integrate multiple audio-visual tasks within dynamic, real-world scenes. However, current works such as event localization, parsing, segmentation and question answering are mostly explored individually, making it challenging to comprehensively understand complex audio-visual scenes and explore inter-task relationships. Hence, we propose \textbf{AV-Unified}, a unified framework that enables joint learning across a wide range of audio-visual scene understanding tasks. AV-Unified standardizes the diverse input-output formats of each task and incorporates a multi-scale spatiotemporal perception network to effectively capture audio-visual associations. Specifically, we unify the inputs and outputs of all supported tasks by converting them into sequences of discrete tokens, establishing a shared representation that allows a single architecture to be trained jointly across heterogeneous varied datasets. Considering the varying temporal granularity of audio-visual events, a multi-scale temporal perception module is designed to capture key cues. Meanwhile, to overcome the lack of auditory supervision in the visual domain, we design a cross-modal guidance-based spatial perception module that models spatial audio-visual associations. Furthermore, task-specific text prompts are employed to enhance the model's adaptability and task-awareness. Extensive experiments on benchmark datasets (e.g., AVE, LLP, MUSIC-AVQA, VGG-SS and AVS) demonstrate the effectiveness of AV-Unified across temporal, spatial, and spatiotemporal tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。