首次评估视觉语言模型在跨文化故事生成中的适应能力,揭示其潜力与局限。
Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation
- 通过扰动多模态文化身份线索,构建新评估框架
- 模型能生成含地域/家族名称等文化特异性词汇,但表现差异大
- 视觉-语义相似性可识别文化差异,但整体理解仍有限
随着视觉语言模型(VLMs)在多元文化场景中广泛应用,确保其文化适应能力对负责任的AI至关重要。现有研究主要关注纯文本模型或对象识别任务中的文化意识,尚未系统评估当文化身份线索同时嵌入文本提示与视觉输入时,VLM在生成任务中的输出适配情况。本文提出首个基于多模态故事生成的文化适应性综合评估,构建新型多模态框架,对5个主流VLM进行测试。分析显示模型具备显著文化适应能力,能生成包含姓名、亲属称谓、地理标识等文化特异性词汇;但存在严重局限:不同架构间表现差异显著,部分模型出现反向文化对齐现象,自动评估指标与人工判断存在偏差。跨模态评估表明,文化差异输出可通过视觉-语义相似性检测(同国籍召回率28.7%,跨国籍0.2%),但视觉-文化理解仍有限。本研究揭示了多模态AI文化适应性的机遇与挑战。代码与数据已公开:https://github.com/ArkaMukherjee0/mmCultural
原文摘要 · Abstract (English)
As Vision-Language Models (VLMs) achieve widespread deployment across diverse cultural contexts, ensuring their cultural competence becomes critical for responsible AI systems. While prior work has evaluated cultural awareness in text-only models and VLM object recognition tasks, no research has systematically assessed how VLMs adapt outputs when cultural identity cues are embedded in both textual prompts and visual inputs during generative tasks. We present the first comprehensive evaluation of VLM cultural competence through multimodal story generation, developing a novel multimodal framework that perturbs cultural identity and evaluates 5 contemporary VLMs on a downstream task: story generation. Our analysis reveals significant cultural adaptation capabilities, with rich culturally-specific vocabulary spanning names, familial terms, and geographic markers. However, we uncover concerning limitations: cultural competence varies dramatically across architectures, some models exhibit inverse cultural alignment, and automated metrics show architectural bias contradicting human assessments. Cross-modal evaluation shows that culturally distinct outputs are indeed detectable through visual-semantic similarity (28.7% within-nationality vs. 0.2% cross-nationality recall), yet visual-cultural understanding remains limited. In essence, we establish the promise and challenges of cultural competence in multimodal AI. We publicly release our codebase and data: https://github.com/ArkaMukherjee0/mmCultural
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。