arXiv:2601.21666cs.AIcs.CV2026-01被引 3

构建首个真实场景音视频理解评测基准,揭示模型在时间定位上的性能差距。

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

  • 基于60小时真实对话视频,设计三类任务评估多模态大模型能力
  • 闭源模型在时间定位上比开源模型高22.6%,跨群体准确率差距达21.4%
  • 提供可复现的评测套件,适合关注公平性与实际应用的研究者

多模态大语言模型(MLLMs)是当前AI研究热点,但多数工作聚焦静态图像理解,对时序音视频数据的处理能力仍缺乏系统评估。为此,我们提出SONIC-O1,一个全面且经人工验证的基准,包含60小时(231段视频)、13个真实对话领域、4,958条标注及人口统计元数据。该基准评估三大能力:开放式摘要、多项选择题回答和带推理依据的时间定位。实验发现,尽管不同模型在多项选择题上差距较小,但最优闭源模型在时间定位上仍比最优开源模型高出22.6%;同时,在不同人口群体间存在最高达21.4%的准确率差异,暴露模型行为中的持续不平等。SONIC-O1为时序对齐与人口多样性鲁棒的多模态理解提供开放评测平台。项目主页、数据集、代码与排行榜均已公开。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are a major focus of recent AI research. However, most prior work focuses on static image understanding, while their ability to process sequential audio-video data remains underexplored. This gap highlights the need for a high-quality benchmark to systematically evaluate MLLM performance in a real-world setting. We introduce SONIC-O1, a comprehensive, fully human-verified benchmark of 60 hours (231 clips) spanning 13 real-world conversational domains with 4,958 annotations and demographic metadata. SONIC-O1 evaluates three capabilities: open-ended summarization, multiple-choice question (MCQ) answering, and temporal localization with supporting rationales (reasoning). Across closed- and open-source models, we find that the MCQ accuracy shows the smallest gap between model families, but the best closed-source model outperforms the best open-source model by 22.6% on temporal localization. We further observe accuracy gaps of up to 21.4% on temporal localization across demographic groups, indicating persistent disparities in model behaviour. SONIC-O1 provides an open evaluation suite for temporally grounded and demographically robust multimodal understanding. SONIC-O1 is publicly available for research: Project page (https://vectorinstitute.github.io/sonic-o1/), Dataset (https://huggingface.co/datasets/vector-institute/sonic-o1), GitHub (https://github.com/vectorinstitute/sonic-o1), Leaderboard (https://huggingface.co/spaces/vector-institute/sonic-o1-leaderboard).

多模态音视频理解评测基准公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。