动态感知音视频关键信息,提升复杂问题理解能力
AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
- 双路径动态采样聚焦问题相关片段
- 模态偏好策略选择性激活关键特征
- 适合需要跨模态推理的多模态任务
音频-视觉问答(AVQA)要求模型有效利用视听模态回答关于音视频场景的复杂问题。现有方法在时间采样灵活性和模态偏好感知方面不足,难以根据问题聚焦关键信息,限制了复杂场景下的推理能力。为此,我们提出AV-Master框架,通过动态建模时间与模态维度,增强从冗余内容中提取关键信息的能力。在时间维度,引入动态自适应聚焦采样机制,逐步关注与问题最相关的音视频片段,有效缓解传统采样中的冗余与片段碎片化问题。在模态维度,提出偏好感知策略,独立建模各模态贡献,实现关键特征的选择性激活。此外,设计双路径对比损失,强化时间与模态维度的一致性与互补性,引导模型学习问题特定的跨模态协作表征。在四个大规模基准测试上,AV-Master显著优于现有方法,尤其在复杂推理任务中表现突出。
原文摘要 · Abstract (English)
Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse questions about audio-visual scenes. However, existing methods lack sufficient flexibility and dynamic adaptability in temporal sampling and modality preference awareness, making it difficult to focus on key information based on the question. This limits their reasoning capability in complex scenarios. To address these challenges, we propose a novel framework named AV-Master. It enhances the model's ability to extract key information from complex audio-visual scenes with substantial redundant content by dynamically modeling both temporal and modality dimensions. In the temporal dimension, we introduce a dynamic adaptive focus sampling mechanism that progressively focuses on audio-visual segments most relevant to the question, effectively mitigating redundancy and segment fragmentation in traditional sampling methods. In the modality dimension, we propose a preference-aware strategy that models each modality's contribution independently, enabling selective activation of critical features. Furthermore, we introduce a dual-path contrastive loss to reinforce consistency and complementarity across temporal and modality dimensions, guiding the model to learn question-specific cross-modal collaborative representations. Experiments on four large-scale benchmarks show that AV-Master significantly outperforms existing methods, especially in complex reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。