arXiv:2504.17821cs.CVcs.CL2025-04ACL被引 3

首个跨文化多语言视频理解评测集,揭示模型在中文语境下的短板。

VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension

  • 构建跨中西欧文化的多语言视频问答数据集
  • 中文相关问题准确率显著低于西方语境,历史类题表现最差
  • 适合评估模型在非英语文化场景下的理解能力

评估多模态AI系统的视频理解能力可有效衡量其认知与推理水平。现有视频评测基准大多仅限于单一语言(如英语),且主要基于西方文化背景。本文提出VideoVista-CulturalLingo,首个旨在弥合文化、语言与领域鸿沟的视频理解评测基准。该数据集涵盖中国、北美和欧洲文化,支持中英文双语提问,并覆盖数百个由人类创作的领域。包含1,389个视频和3,134组问答对,我们评估了24个近期开源或商用视频大模型。实验发现:1)现有模型在中文语境问题上表现更差,尤其在涉及中国历史的问题上;2)当前开源模型在时间理解任务(事件定位)上仍有局限,最高得分仅45.2%;3)主流模型在通用科学问题上表现良好,但开源模型在数学类问题上能力较弱。

原文摘要 · Abstract (English)

Assessing the video comprehension capabilities of multimodal AI systems can effectively measure their understanding and reasoning abilities. Most video evaluation benchmarks are limited to a single language, typically English, and predominantly feature videos rooted in Western cultural contexts. In this paper, we present VideoVista-CulturalLingo, the first video evaluation benchmark designed to bridge cultural, linguistic, and domain divide in video comprehension. Our work differs from existing benchmarks in the following ways: 1) Cultural diversity, incorporating cultures from China, North America, and Europe; 2) Multi-linguistics, with questions presented in Chinese and English-two of the most widely spoken languages; and 3) Broad domain, featuring videos sourced from hundreds of human-created domains. VideoVista-CulturalLingo contains 1,389 videos and 3,134 QA pairs, and we have evaluated 24 recent open-source or proprietary video large models. From the experiment results, we observe that: 1) Existing models perform worse on Chinese-centric questions than Western-centric ones, particularly those related to Chinese history; 2) Current open-source models still exhibit limitations in temporal understanding, especially in the Event Localization task, achieving a maximum score of only 45.2%; 3) Mainstream models demonstrate strong performance in general scientific questions, while open-source models demonstrate weak performance in mathematics.

视频理解多语言跨文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。