arXiv:2609.06831cs.AIcs.CL2026-09

构建跨文化视觉规范理解基准,让AI学会看懂不同文化中的行为含义。

NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures

论文配图:NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures
图 1 · 摘自论文原文
  • 设计6536张对比图像对,测试模型对文化行为规范的判断能力
  • 顶尖多模态模型在跨文化行为识别上准确率不足27%
  • 提供带解释的6.4万张训练数据,助力模型理解视觉与文化关联

AI系统在全球应用广泛,但难以满足多元文化群体需求。现有研究多聚焦文本或视觉物品识别(如食物、服饰),对通过视觉行为推断本地社会规范的能力——即视觉规范理解——尚未考察。本文提出NormViz-Bench,一个高质量、人工验证的基准,包含3,268对对比图像(共6,536张),覆盖16个国家。每对图像仅在文化相关行为(如物体、属性、空间关系、动作)上存在差异,影响图像解读。每张图标注为符合、违反或无关当地社会规范,配对评估要求两图均正确分类,防止模型依赖表面视觉线索。即使最强模型Gemini 3.0 Flash和Qwen2.5 VL 7B,准确率也仅为26.6%和21.6%,尤其在识别违规与文化中性行为时表现差。为此,我们构建NormViz-Train,包含64,000张配解释的图像。尽管绝对性能仍低于30%,但在该数据集微调后,Qwen3-VL 4B和8B的配对准确率相对提升达125%和36%,表明模型具备学习视觉与文化意义关联的潜力。综上,NormViz-Bench与NormViz-Train确立了视觉规范理解作为多模态AI的关键挑战与发展方向。

原文摘要 · Abstract (English)

AI systems are used worldwide, but they struggle to serve the needs of culturally diverse populations. Prior work on cultural understanding evaluates AI systems on text-only settings or on visual artifact recognition (e.g. foods, clothing). The ability to reason about visually observable behaviors through local social norms, which we call visual norm understanding, remains unexamined. We introduce NormViz-Bench, a high quality, human-validated benchmark of 3,268 contrastive image pairs (6,536 images) spanning 16 countries. Each pair varies only in the culturally relevant behavior (e.g., objects, attributes, spatial relations, and actions) that alters how each image is interpreted. Each image is labeled as conforming to, violating, or irrelevant to local social norms, and pair-level evaluation requires both images to be correctly classified, thereby preventing reliance on superficial visual shortcuts. Even the strongest VLMs, Gemini 3.0 Flash and Qwen2.5 VL 7B, succeed on only 26.6% and 21.6% of pairs, struggling most with identifying violating and culturally benign visual behaviors. Towards bridging this, we introduce NormViz-Train, a training dataset of 64k images paired with explanations. Though absolute performance remains low (<30%), finetuning on NormViz-Train improves pair accuracy relatively by up to 125% and 36% Qwen3-VL 4B and 8B respectively, showing a path forward to teach models to connect visual perception to cultural significance. Together, NormViz-Bench and NormViz-Train establish visual norm understanding as a challenging and consequential frontier for multimodal AI.

多模态推理文化理解视觉规范基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。