提出新方法融合问题引导的音视频时空频域特征,提升多模态问答性能。
Query-Guided Spatial-Temporal-Frequency Interaction for Music Audio-Visual Question Answering
- 用问题引导音视频的时空频域交互,增强跨模态理解
- 在多个基准上超越现有音频、视觉及视频问答方法
- 适合研究多模态推理与视听问答的学者参考
音频-视觉问答(AVQA)是一项需要联合分析视频中音频、视觉和文本信息以回答自然语言问题的挑战性多模态任务。尽管近期视频问答进展显著,现有AVQA方法仍主要依赖预训练模型提取物体级和运动级视觉表征,而将音频视为辅助信息,文本问题仅在推理后期被整合,贡献有限。为此,本文提出一种新型查询引导的时空频域交互(QSTar)方法,通过引入问题引导线索,并结合音频信号的频率特性,协同空间与时间感知,增强音频-视觉理解。此外,受提示机制启发,设计了查询上下文推理(QCR)模块,使模型更精准聚焦于语义相关的音视频特征。在多个AVQA基准上的大量实验表明,所提方法显著优于现有音频问答、视觉问答、视频问答及多模态问答方法。代码与预训练模型将在论文发表后公开。
原文摘要 · Abstract (English)
Audio--Visual Question Answering (AVQA) is a challenging multimodal task that requires jointly reasoning over audio, visual, and textual information in a given video to answer natural language questions. Inspired by recent advances in Video QA, many existing AVQA approaches primarily focus on visual information processing, leveraging pre-trained models to extract object-level and motion-level representations. However, in those methods, the audio input is primarily treated as complementary to video analysis, and the textual question information contributes minimally to audio--visual understanding, as it is typically integrated only in the final stages of reasoning. To address these limitations, we propose a novel Query-guided Spatial--Temporal--Frequency (QSTar) interaction method, which effectively incorporates question-guided clues and exploits the distinctive frequency-domain characteristics of audio signals, alongside spatial and temporal perception, to enhance audio--visual understanding. Furthermore, we introduce a Query Context Reasoning (QCR) block inspired by prompting, which guides the model to focus more precisely on semantically relevant audio and visual features. Extensive experiments conducted on several AVQA benchmarks demonstrate the effectiveness of our proposed method, achieving significant performance improvements over existing Audio QA, Visual QA, Video QA, and AVQA approaches. The code and pretrained models will be released after publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。