测试视觉语言模型对火文化图像的理解,发现其常误判非西方节日和紧急场景。
Seeing Symbols, Missing Cultures: Probing Vision-Language Models' Reasoning on Fire Imagery and Cultural Meaning
- 用分类与解释分析结合的诊断框架评估模型对火意象的文化理解
- 模型能识别西方节日但误判非西方传统和紧急事件,错误率超40%
- 适合关注多模态公平性与可解释性的研究者阅读
视觉语言模型(VLMs)看似具备文化理解能力,实则依赖表面模式匹配而非真实文化认知。本文提出一种诊断框架,通过分类与解释分析相结合的方式,探测VLM在火主题文化图像上的推理表现。在测试多个模型时,发现系统性偏差:模型能准确识别显著的西方节日,但在非西方传统和紧急场景中表现不佳,常给出模糊标签或错误地将紧急情况误判为庆祝活动。这些失败揭示了符号化捷径的风险,强调需超越准确率指标,在多模态系统中引入文化层面的评估机制,以实现更可解释、更公平的AI应用。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) often appear culturally competent but rely on superficial pattern matching rather than genuine cultural understanding. We introduce a diagnostic framework to probe VLM reasoning on fire-themed cultural imagery through both classification and explanation analysis. Testing multiple models on Western festivals, non-Western traditions, and emergency scenes reveals systematic biases: models correctly identify prominent Western festivals but struggle with underrepresented cultural events, frequently offering vague labels or dangerously misclassifying emergencies as celebrations. These failures expose the risks of symbolic shortcuts and highlight the need for cultural evaluation beyond accuracy metrics to ensure interpretable and fair multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。