arXiv:2509.19274cs.CLcs.MM2025-09EMNLP被引 18

首个面向印度文化的多模态多语言评测基准,测试AI对本土文化的理解能力。

DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture

  • 构建覆盖15种语言、64000组图文对的印度文化多模态数据集
  • 发现主流模型在低资源语言和冷门传统上推理能力严重不足
  • 适合研究跨文化AI、本地化大模型与包容性技术的学者使用

我们提出DRISHTIKON,首个专注于印度文化的多模态多语言评测基准,旨在评估生成式AI系统对本土文化的理解能力。与现有通用或全球性基准不同,该数据集涵盖印度全部15个语言、所有州与联邦属地,包含超过64,000组对齐的图文对,细致覆盖节庆、服饰、饮食、艺术形式及历史遗产等丰富文化主题。我们在零样本与思维链(chain-of-thought)设置下,评估了多种视觉-语言模型(VLMs),包括开源小规模与大规模模型、专有系统、专注推理的VLMs以及针对印地语系优化的模型。结果揭示当前模型在基于文化背景的多模态输入上存在显著局限,尤其在低资源语言和非主流传统中表现不佳。DRISHTIKON填补了包容性AI研究的关键空白,为推进具备文化敏感性的多模态语言技术提供可靠测试平台。

原文摘要 · Abstract (English)

We introduce DRISHTIKON, a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture, designed to evaluate the cultural understanding of generative AI systems. Unlike existing benchmarks with a generic or global scope, DRISHTIKON offers deep, fine-grained coverage across India's diverse regions, spanning 15 languages, covering all states and union territories, and incorporating over 64,000 aligned text-image pairs. The dataset captures rich cultural themes including festivals, attire, cuisines, art forms, and historical heritage amongst many more. We evaluate a wide range of vision-language models (VLMs), including open-source small and large models, proprietary systems, reasoning-specialized VLMs, and Indic-focused models, across zero-shot and chain-of-thought settings. Our results expose key limitations in current models' ability to reason over culturally grounded, multimodal inputs, particularly for low-resource languages and less-documented traditions. DRISHTIKON fills a vital gap in inclusive AI research, offering a robust testbed to advance culturally aware, multimodally competent language technologies.

文化理解多模态多语言AI评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。