arXiv:2510.08608cs.CLcs.AI2025-10被引 3

构建首个跨模态亚洲文化评估框架,测试大模型在多元文化下的理解能力。

MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation

  • 设计多语言多模态对齐的评测集,覆盖8国10语种2.7万题
  • 超79%题目需文化背景支撑的多步推理,避免机械记忆
  • 提出五维评估体系,可检测模型跨语言跨模态偏差

大型语言模型在全球广泛应用,但其多模态理解和推理能力在非西方、低资源地区常显著下降。本文提出MMA-ASIA框架,聚焦亚洲文化情境,评估大模型的文化敏感性。该框架包含一个由人工标注的多语言、多模态对齐的多项选择基准,覆盖8个亚洲国家和10种语言,共27,000道题目;其中超过79%的问题需要基于文化背景的多步推理,超越简单记忆。据我们所知,这是首个在文本、图像(视觉问答)和语音三个模态输入层面实现对齐的数据集,支持直接的跨模态迁移测试。基于此基准,我们提出五维评估协议:(i) 国家间文化感知差异,(ii) 跨语言一致性,(iii) 跨模态一致性,(iv) 文化知识泛化能力,(v) 知识锚定有效性。为确保严谨评估,引入文化意识锚定验证模块,检测是否存在“捷径学习”——即答案是否依赖于必要文化知识。通过模型对比分析、注意力追踪及创新的视觉缺失前缀重播(VPR)方法,揭示模型在不同语言和模态间的差异成因,为构建更具文化可靠性的多模态大模型提供可操作洞见。

原文摘要 · Abstract (English)

Large language models (LLMs) are now used worldwide, yet their multimodal understanding and reasoning often degrade outside Western, high-resource settings. We propose MMA-ASIA, a comprehensive framework to evaluate LLMs' cultural awareness with a focus on Asian contexts. MMA-ASIA centers on a human-curated, multilingual, and multimodally aligned multiple-choice benchmark covering 8 Asian countries and 10 languages, comprising 27,000 questions; over 79 percent require multi-step reasoning grounded in cultural context, moving beyond simple memorization. To our knowledge, this is the first dataset aligned at the input level across three modalities: text, image (visual question answering), and speech. This enables direct tests of cross-modal transfer. Building on this benchmark, we propose a five-dimensional evaluation protocol that measures: (i) cultural-awareness disparities across countries, (ii) cross-lingual consistency, (iii) cross-modal consistency, (iv) cultural knowledge generalization, and (v) grounding validity. To ensure rigorous assessment, a Cultural Awareness Grounding Validation Module detects "shortcut learning" by checking whether the requisite cultural knowledge supports correct answers. Finally, through comparative model analysis, attention tracing, and an innovative Vision-ablated Prefix Replay (VPR) method, we probe why models diverge across languages and modalities, offering actionable insights for building culturally reliable multimodal LLMs.

多模态评估文化敏感性跨语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。