为视觉语言模型设计文化安全评估框架,发现主流模型在文化合规上仍严重不足。
Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
- 构建跨文化基准CROSS,结合图像与多语言上下文测试文化理解能力
- 21个主流模型中最佳仅61.79%文化意识,37.73%行为合规,差距明显
- 通过强化学习和指令微调可显著提升文化安全,且不影响通用多模态能力
大型视觉语言模型(LVLMs)在旅游助手等全球化应用中日益普及,但其生成符合文化规范回应的能力尚未充分探索。现有多模态安全基准主要关注物理安全,忽视了由文化规范缺失导致的象征性伤害。为此,我们提出CROSS基准,包含来自16个国家、14种语言、三个日常场景的1,284个图文关联问题,文化违规仅在上下文图像理解下显现。我们进一步提出基于跨文化理论的CROSS-Eval框架,衡量文化意识、规范教育、合规性与助人度四项指标。评估21个领先LVLMs显示:最佳模型文化意识仅达61.79%,合规性仅37.73%。部分开源模型虽接近GPT-4o表现,但仍显著落后于闭源模型。增加推理能力有助于改善文化对齐,但无法根本解决。为此我们提出两种增强策略:基于文化语境的开放式监督微调,以及对比响应对的偏好调优。这两种方法使GPT-4o的文化意识提升60.14%,合规性提升55.2%,同时在通用多模态理解基准上性能下降极小。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) are increasingly deployed in globally distributed applications, such as tourism assistants, yet their ability to produce culturally appropriate responses remains underexplored. Existing multimodal safety benchmarks primarily focus on physical safety and overlook violations rooted in cultural norms, which can result in symbolic harm. To address this gap, we introduce CROSS, a benchmark designed to assess the cultural safety reasoning capabilities of LVLMs. CROSS includes 1,284 multilingual visually grounded queries from 16 countries, three everyday domains, and 14 languages, where cultural norm violations emerge only when images are interpreted in context. We propose CROSS-Eval, an intercultural theory-based framework that measures four key dimensions: cultural awareness, norm education, compliance, and helpfulness. Using this framework, we evaluate 21 leading LVLMs, including mixture-of-experts models and reasoning models. Results reveal significant cultural safety gaps: the best-performing model achieves only 61.79% in awareness and 37.73% in compliance. While some open-source models reach GPT-4o-level performance, they still fall notably short of proprietary models. Our results further show that increasing reasoning capacity improves cultural alignment but does not fully resolve the issue. To improve model performance, we develop two enhancement strategies: supervised fine-tuning with culturally grounded, open-ended data and preference tuning with contrastive response pairs that highlight safe versus unsafe behaviors. These methods substantially improve GPT-4o's cultural awareness (+60.14%) and compliance (+55.2%), while preserving general multimodal capabilities with minimal performance reduction on general multimodal understanding benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。