arXiv:2512.01419cs.CVcs.AI2025-12被引 2

构建东盟文化理解测评集,揭示视觉语言模型的跨文化偏差

Rice-VL: Evaluating Vision-Language Models for Cultural Understanding Across ASEAN Countries

  • 设计覆盖11国2.8万+样本的多模态评测集
  • 发现低资源国家与抽象文化领域性能显著下降
  • 适合关注AI公平性与跨文化AI研究者使用

视觉语言模型在多模态任务中表现优异,但常存在西方中心偏见,限制其在东南亚(SEA)等文化多元地区的效果。为此,我们提出RICE-VL,一个评估视觉语言模型在11个东盟国家文化理解能力的新基准。该数据集包含超过28,000条人工标注的视觉问答(VQA)样本,涵盖是非判断、填空和开放式问题,并有1,000对图像边界框用于视觉定位,由14个子文化类别中的文化专家标注。我们提出SEA-LAVE,作为LAVE度量的扩展,评估文本准确性、文化契合度和国家识别能力。对六种开源与闭源模型的评估显示,在低资源国家及抽象文化领域存在显著性能差距。视觉定位任务测试模型在复杂场景中定位文化关键元素的能力,检验空间与上下文准确性。RICE-VL揭示了视觉语言模型在文化理解上的局限,强调需推动更具包容性的模型开发以更好服务全球多样性人群。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) excel in multimodal tasks but often exhibit Western-centric biases, limiting their effectiveness in culturally diverse regions like Southeast Asia (SEA). To address this, we introduce RICE-VL, a novel benchmark evaluating VLM cultural understanding across 11 ASEAN countries. RICE-VL includes over 28,000 human-curated Visual Question Answering (VQA) samples -- covering True or False, Fill-in-the-Blank, and open-ended formats -- and 1,000 image-bounding box pairs for Visual Grounding, annotated by culturally informed experts across 14 sub-ground categories. We propose SEA-LAVE, an extension of the LAVE metric, assessing textual accuracy, cultural alignment, and country identification. Evaluations of six open- and closed-source VLMs reveal significant performance gaps in low-resource countries and abstract cultural domains. The Visual Grounding task tests models' ability to localize culturally significant elements in complex scenes, probing spatial and contextual accuracy. RICE-VL exposes limitations in VLMs' cultural comprehension and highlights the need for inclusive model development to better serve diverse global populations.

视觉语言模型文化理解多模态评测东盟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。