arXiv:2505.21979cs.CL2025-05EMNLP被引 7

构建首个面向阿拉伯文化的多模态数据集,助力模型理解文化差异

Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset

  • 通过智能代理与37位阿拉伯地区专家协作标注,构建超30万条多模态数据
  • 新基准测试显示,以推理对齐的模型比传统扩展方法文化理解更强
  • 适合研究多模态模型文化偏见、跨文化理解或阿拉伯语AI应用者使用

主流大规模视觉语言模型(LVLM)普遍存在文化偏见,亟需多样化多模态数据集。为此,我们提出PEARL,一个大规模阿拉伯语多模态数据集及评估基准,专为文化理解设计。该数据集通过先进智能体工作流与来自阿拉伯世界37位标注者的深度人工协作构建,包含超过30.9万条多模态样本,覆盖10个具有文化意义的领域,涵盖所有阿拉伯国家。我们进一步提供两个稳健的评估基准(PEARL和PEARL-LITE),以及一个专门用于检测细微文化差异的子集(PEARL-X)。对前沿开源与闭源LVLM的全面评估表明,以推理为中心的指令对齐显著提升模型的文化基底能力,优于传统规模扩展方法。所有数据集与基准均公开可用。

原文摘要 · Abstract (English)

Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. Constructed through advanced agentic workflows and extensive human-in-the-loop annotations by 37 annotators from across the Arab world, PEARL comprises over 309K multimodal examples spanning ten culturally significant domains covering all Arab countries. We further provide two robust evaluation benchmarks (PEARL and PEARL-LITE) along with a specialized subset (PEARL-X) explicitly developed to assess nuanced cultural variations. Comprehensive evaluations on state-of-the-art open and proprietary LVLMs demonstrate that reasoning-centric instruction alignment substantially improves models' cultural grounding compared to conventional scaling methods. PEARL establishes a foundational resource for advancing culturally-informed multimodal modeling research. All datasets and benchmarks are publicly available.

多模态文化理解阿拉伯语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。