arXiv:2608.23065cs.CVcs.AI2026-08中稿 · EMNLP

评测视频文化理解的东南亚文化基准,区分命名、识别和定位三能力。

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

论文配图:Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia
图 1 · 摘自论文原文
  • 分三阶段测试文化概念的命名、视觉识别和时间定位能力。
  • 最强模型全对率低于30%,且三种能力不完全连锁。
  • 揭示音频在非拉丁文字国家常起干扰作用,适合跨文化研究者。

视频中的文化理解不仅关乎可见内容,更需把握文化概念的象征与时间意义。我们将其分解为三项能力:命名概念所象征的意义、在视频中视觉识别该概念、定位其子事件的时间段。现有视频文化基准多仅测试可见性,将三能力合并评分,掩盖瓶颈。为此提出东南亚文化时刻基准(CMB):涵盖东南亚七国306个专家标注的概念,分为五大类。每概念通过三阶段评估,分别对应一项能力。第一阶段(S1)根据描述从四个候选名称中选择;第二阶段(S2)从四个候选视频片段中选择;第三阶段(S3)在另一视频中预测该片段的起止时间。为确保各阶段聚焦单一能力,采用语义相似干扰项(S1、S2)、未标注视频片段(S2)及不同示例视频的自由形式定位(S3)。六种视觉-语言模型的表现显示失败模式因能力与模态而异:即使最强闭源模型在三阶段全对时得分仍低于30%;三能力不完全递进:正确命名可帮助一半模型识别视频内容,但识别对定位时间影响甚微;音频对部分概念具互补性,对其他则冗余或干扰,尤其在非拉丁字母国家更易造成干扰;去除音频与字幕后,游戏与音乐类任务受损最严重。14名评审的人工评估表明,即使专家在邻国概念上也表现不及随机水平,说明CMB依赖国别文化知识。该基准可作为诊断工具,精准定位失败归因于特定能力或模态。

原文摘要 · Abstract (English)

Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this into three abilities: naming what a concept symbolizes, visually recognizing it on video, and locating its sub-events in time. Existing video-cultural benchmarks tend to test what is seen, collapsing these three abilities into a single score that hides the bottleneck. We introduce the Cultural Moment Benchmark (CMB): 306 expert-curated concepts from seven countries in Southeast Asia across five categories. We evaluate each concept through three stages, one per ability. Given a description, Stage 1 (S1) selects from four candidate concept names, Stage 2 (S2) selects from four candidate video moments, and Stage 3 (S3) predicts the start and end times of the moment in a video. To keep each stage focused on a distinct ability, we use three design choices: semantic-similarity distractors (S1, S2), unlabeled video moments (S2), and free-form localization on a different example video (S3). Across six vision-language models, failure modes vary by ability and modality. i) Even the strongest closed-source models score below 30% when all three stages must be correct; ii) The three abilities do not fully cascade: naming a concept correctly helps half the models recognize it on video, but recognizing it has little effect on locating the sub-event in time; iii) Audio is complementary, redundant, or distracting depending on the concept, more often distracting in non-Latin-script countries; removing both audio and subtitles hurts Games and Music the most. Our 14-rater human study shows that even Expert raters score below chance on concepts from a neighboring country, indicating that CMB requires country-specific cultural knowledge. CMB acts as a diagnostic harness, attributing failures to a specific ability or modality.

文化理解视频评估多模态东南亚

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。