构建多文化多语言视频推理基准,挑战主流模型的文化理解能力
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
- 基于18个全球地区真实文化视频,人工标注跨语言复杂推理题
- 顶尖视频大模型在该基准上表现远低于人类,错误主要源于文化视觉识别
- 提出基于推理轨迹的图结构迭代纠错方法,助力细粒度错误定位
近期视频模型在长视频理解方面取得显著进展,但现有评估基准多以西方内容为主、以英语为单一语言,造成明显偏见。为此,我们提出MINERVA-Cultural,一个面向多元文化和多语言视频推理的高难度基准。该数据集包含来自18个全球区域的真实文化视频,所有问题、答案及多步推理过程均由母语者人工构建,避免自动翻译偏差。真正掌握该基准需对视觉文化语境有深度理解。此外,我们利用推理轨迹构建证据图,并提出一种新颖的迭代策略,用于识别推理中的细粒度错误。评估显示,当前最先进视频大模型表现远低于人类水平,主要错误源于对文化元素的视觉感知不足。MINERVA-Cultural将公开发布于https://github.com/google-deepmind/neptune?tab=readme-ov-file#minerva-cultural。
原文摘要 · Abstract (English)
Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data and English as the dominant language, introducing significant biases in evaluation. To address this, we introduce MINERVA-Cultural, a challenging benchmark for multicultural and multilingual video reasoning. MINERVA-Cultural comprises high-quality, entirely human-generated annotations from diverse, region-specific cultural videos across 18 global locales. Unlike prior work that relies on automatic translations, MINERVA-Cultural provides complex questions, answers, and multi-step reasoning steps, all crafted in native languages. Making progress on MINERVA-Cultural requires a deeply situated understanding of visual cultural context. Furthermore, we leverage MINERVA-Cultural's reasoning traces to construct evidence-based graphs and propose a novel iterative strategy using these graphs to identify fine-grained errors in reasoning. Our evaluations reveal that SoTA Video-LLMs struggle significantly, performing substantially below human-level accuracy, with errors primarily stemming from the visual perception of cultural elements. MINERVA-Cultural will be publicly available under https://github.com/google-deepmind/neptune?tab=readme-ov-file\#minerva-cultural
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。