arXiv:2510.17722cs.CVcs.AI2025-10ACL被引 7

评测多模态大模型在多轮视频对话中的理解能力

MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues

  • 构建多轮对话视频理解评测集,覆盖6项核心能力
  • 包含1000组跨领域多轮对话,贴近真实应用场景
  • 适合研究多模态交互与视频理解的学者使用

多模态大语言模型(MLLMs)在视觉理解方面取得了显著进展,但现有评测基准仍局限于单轮问答,忽视了现实场景中多轮对话的复杂性。为此,我们提出MT-Video-Bench,一个全面评估MLLM在多轮对话中视频理解能力的基准。该基准重点评估感知与交互相关的6项核心能力,涵盖来自多个领域的1000组精心设计的多轮对话,紧密对齐真实应用,如互动体育分析和多轮视频智能辅导。通过该基准,我们对多种先进的开源与闭源MLLM进行了广泛评估,揭示了其在处理多轮视频对话时存在的显著性能差异与局限性。该评测集将公开发布,以推动后续研究。

原文摘要 · Abstract (English)

The recent development of Multimodal Large Language Models (MLLMs) has significantly advanced AI's ability to understand visual modalities. However, existing evaluation benchmarks remain limited to single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios. To bridge this gap, we introduce MT-Video-Bench, a holistic video understanding benchmark for evaluating MLLMs in multi-turn dialogues. Specifically, our MT-Video-Bench mainly assesses 6 core competencies that focus on perceptivity and interactivity, encompassing 1,000 meticulously curated multi-turn dialogues from diverse domains. These capabilities are rigorously aligned with real-world applications, such as interactive sports analysis and multi-turn video-based intelligent tutoring. With MT-Video-Bench, we extensively evaluate various state-of-the-art open-source and closed-source MLLMs, revealing their significant performance discrepancies and limitations in handling multi-turn video dialogues. The benchmark will be publicly available to foster future research.

多模态视频理解对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。