arXiv:2511.18399cs.CV2025-11被引 3

首个专用于中文视频问答的多模态大模型评测基准

ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering

  • 构建涵盖8大类12子类的中文视频理解任务集
  • 现有多模态模型在该基准上最高得分仅77.9%
  • 适合评估中文语境下视觉语言模型的跨模态能力

本文提出ChineseVideoBench,首个针对中文视频问答任务的多模态大模型评测基准。该基准包含8个主类别和12个子类别,覆盖需深度视频理解与汉语语言文化认知的任务,旨在评估先进多模态大模型在复杂中文视频内容上的表现。实证评估显示,当前模型面临严峻挑战:Gemini 2.5 Pro取得最高综合得分77.9%,而InternVL-38B是表现最出色的开源模型。

原文摘要 · Abstract (English)

This paper introduces ChineseVideoBench, a pioneering benchmark specifically designed for evaluating Multimodal Large Language Models (MLLMs) in Chinese Video Question Answering. The growing demand for sophisticated video analysis capabilities highlights the critical need for comprehensive, culturally-aware evaluation frameworks. ChineseVideoBench addresses this gap by providing a robust dataset and tailored evaluation metrics, enabling rigorous assessment of state-of-the-art MLLMs on complex Chinese video content. Specifically, ChineseVideoBench comprises 8 main classes and 12 sub-classes, encompassing tasks that demand both deep video understanding and nuanced Chinese linguistic and cultural awareness. Our empirical evaluations reveal that ChineseVideoBench presents a significant challenge to current MLLMs. Among the models assessed, Gemini 2.5 Pro achieves the highest performance with an overall score of 77.9%, while InternVL-38B emerges as the most competitive open-source model.

视频问答多模态中文评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。