arXiv:2607.10299cs.LG2026-07被引 1

构建首个解耦音视频描述数据集,提升长时多模态理解能力

Empowering Long-form Omni-modal Understanding with Robust Audio Perception

论文配图:Empowering Long-form Omni-modal Understanding with Robust Audio Perception
图 1 · 摘自论文原文
  • 用自动流水线生成视觉、音频、音视频联合三类标注,分离模态语义
  • 在多个下游任务中实现显著性能提升,音视频联合推理能力增强
  • 适合研究多模态融合、语音感知或视频理解的开发者使用

近期大规模多模态模型在视觉-语言任务上取得显著进展,但全面的跨模态理解仍受限于缺乏富含明确对齐听觉线索的数据集。为此,我们提出 AVDC(Audio-Visual Decoupled Captions)——一个大规模数据集,旨在解耦视觉与听觉语义。具体而言,我们设计自动化流程,利用现成模型为视频生成三类标注:仅视觉(V)、仅音频(A)和音视频联合(AV)。该解耦结构显式捕捉各模态特有细节及复杂跨模态交互。基于此,我们构建了 AVDC-QA-CoT,一个链式思维增强的问答数据集,以促进音视频推理能力。为充分挖掘资源,采用两阶段训练策略:先在 AVDC 上进行全模态描述生成预训练,再在 AVDC-QA-CoT 上进行指令微调。在视频描述、音频主导分析及跨模态基准等多个下游任务上的广泛实验表明,该方法持续且显著提升性能,验证了所提数据集与训练策略在推进跨模态感知方面的有效性。代码与数据集已公开于 https://radiant0726.github.io/AVDC-web/。

原文摘要 · Abstract (English)

Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues. To bridgethis gap, we present AVDC (Audio-Visual Decoupled Captions), a large-scaledataset designed to disentangle visual and auditory semantics. Specifi-cally, we propose an automated pipeline that leverages off-the-shelf mod-els to annotate videos with tripartite captions: visual-only (V), audio-only (A), and joint audio-visual (AV). This decoupled structure explic-itly captures both modality-specific nuances and complex cross-modalinteractions. Building upon this, we introduce AVDC-QA-CoT, a Chain-of-Thought augmented question-answering dataset to foster audio-visualreasoning. To fully exploit these resources, we employ a two-stage train-ing paradigm: omni-modal caption generation pre-training on AVDC, fol-lowed by instruction tuning on AVDC-QA-CoT. Extensive experiments acrossdiverse downstream tasks, spanning video captioning, audio-centric anal-ysis, and omni-modal benchmarks, demonstrate consistent and signifi-cant performance gains, showing the efficacy of our proposed datasetsand training strategy in advancing omni-modal perception. Code anddataset are related on https://radiant0726.github.io/AVDC-web/.

多模态音频理解数据集视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。