构建跨模态时间对齐的音视频问答基准,揭示大模型同步处理能力短板。
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
- 设计半自动标注流程,实现音视频时间对齐与一致性校验
- 在684段真实视频上测试24个模型,发现多数模型在对齐任务中表现不佳
- 提供无需训练的模块化基线,便于诊断模型对齐能力
近期多模态大语言模型在视觉和音频基准上表现良好,但其同步处理跨模态信息的能力仍缺乏探索。本文提出Daily-Omni,一个包含684个真实世界视频和1,197个问题的多选音视频问答基准,涵盖6类任务,明确要求跨模态时间推理。为支持可扩展的基准构建,开发了半自动的标注、跨模态一致性优化、时间对齐提取及纯文本泄露过滤流程,并辅以人工验证。进一步提供诊断评估套件,对24个基础模型在37种模型-模态组合(音频+视频 / 音频仅用 / 视频仅用 / 文本仅用)下进行广泛评估。最后引入一个无需训练的模块化诊断基线,通过组合现成单模态模型来诊断性能并揭示显式时间对齐信号的影响。结果表明,许多端到端多模态模型在依赖对齐的问题上仍表现欠佳,说明稳健的跨模态时间对齐仍是重要开放挑战。
原文摘要 · Abstract (English)
Recent Multimodal Large Language Models (MLLMs) achieve promising performance on visual and audio benchmarks independently. However, the ability of these models to process cross-modal information synchronously remains largely unexplored. We introduce Daily-Omni, a multiple-choice Audio-Visual QA benchmark featuring 684 real-world videos and 1,197 questions spanning 6 task families that explicitly require cross-modal temporal reasoning. To support scalable benchmark construction, we develop a semi-automatic pipeline for annotation, cross-modal consistency refinement, temporal alignment elicitation, and text-only leakage filtering, followed by human verification. We further provide a diagnostic evaluation suite and extensively evaluate 24 foundation models under 37 model--modality settings (Audio+Video / Audio-only / Video-only / Text-only). Finally, we include a training-free modular diagnostic baseline that composes off-the-shelf unimodal models to serve as a diagnostic baseline and to illustrate how explicit temporal alignment signals affect performance. Results indicate that many end-to-end MLLMs still struggle on alignment-critical questions, suggesting that robust cross-modal temporal alignment remains an important open challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。