arXiv:2607.16107eess.AScs.CV2026-07被引 1

开源模型可理解长视频中的音画信息并推理复杂事件。

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

论文配图:Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
图 1 · 摘自论文原文
  • 用三阶段课程训练,从短时感知逐步过渡到长时多事件推理。
  • 在15+评测中超越同类开源模型,部分任务优于大体量闭源模型。
  • 适合需要跨模态长视频分析的研究者和开发者使用。

我们提出 Audio-Visual Flamingo (AV-Flamingo),一个完全开源的先进音画大语言模型(AV-LLM),用于联合理解与推理音频、图像及长视频内容。不同于以往主要针对短片段的音画大模型,AV-Flamingo专为理解与推理长而复杂的现实世界音视频内容设计。为此,我们做出三项关键贡献:(i) 构建包含约700万条标注与问答数据的真实世界视频数据集 Audio-Visual-Skills,强调时间、组合与跨模态音画推理;(ii) 提出一种三阶段渐进式训练策略,从短时感知逐步过渡至长时多事件推理;(iii) 设计时序音画交错思维链(Temporal Audio-Visual Interleaved Chain-of-Thought)推理框架,将中间推理步骤显式锚定在长音视频流的时间戳上,提升时间对齐与可解释性。在15个以上音画、全模态、音频与视觉基准上的大量实验表明,AV-Flamingo显著优于同等规模的开源模型,并在多项任务上超越甚至媲美更大规模的开源与闭源模型,尤其在长且复杂的现实音视频理解与推理任务中表现突出。除基准性能外,该模型展现出强现实应用价值,能良好迁移至未见任务,凸显其鲁棒性与泛化能力。

原文摘要 · Abstract (English)

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.

音画融合长视频理解大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。