构建跨文化知识注入评估基准,提升多模态大模型文化适应性。
CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs

- 提出在不改变原行为前提下,向模型注入特定文化知识的新方法。
- 基准涵盖9800个跨文化图文案例,覆盖49个文化场景。
- 发现现有方法难以兼顾文化适配与非目标文化行为的保持。
多模态大语言模型(MLLMs)主要基于英语数据训练,在跨文化场景中常生成文化不当或不一致的回答。为缓解此问题,我们提出跨文化知识注入任务,旨在适配特定文化背景的同时保留其在其他文化中的原始行为。为此,我们构建了CrossCult-KIBench,一个全面的评估基准,用于衡量知识注入的有效性及其对非目标文化行为的意外影响。该基准包含9800个图像-语境关联的案例,覆盖英、中、阿三个语言-文化群体的49个文化相关视觉场景,支持单次注入与序列注入两种评估设置。我们还提出了基线方法记忆条件知识注入(MCKI),通过冻结的MLLM表示从外部记忆中检索相关文化知识,并在适用时将其作为条件提示前置。在CrossCult-KIBench上的大量实验表明,当前方法难以在有效文化适配与行为保留之间取得平衡,凸显了发展更具文化感知力的MLLM的关键挑战。本工作强调了开发更文化适应且负责任的MLLM的重要研究方向。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs), trained primarily on English-centric data, frequently generate culturally inappropriate or misaligned responses in cross-cultural settings. To mitigate this, we introduce the task of cross-cultural knowledge insertion, which focuses on adapting models to specific cultural contexts while preserving their original behavior in other cultures. To facilitate research in this area, we introduce CrossCult-KIBench, a comprehensive evaluation benchmark for assessing both the effectiveness of knowledge insertion and its unintended side effects on non-target cultures. The benchmark includes 9,800 image-grounded cases covering 49 culturally relevant visual scenarios across English, Chinese, and Arabic language-culture groups. It supports evaluation in both single-insert and sequential-insert settings. We also propose Memory-Conditioned Knowledge Insertion (MCKI) as a baseline method. MCKI retrieves relevant cultural knowledge from an external memory using frozen MLLM representations, prepending matched entries as conditional prompts when applicable. Extensive experiments on CrossCult-KIBench reveal that current approaches struggle to balance effective cultural adaptation with behavioral preservation, highlighting a key challenge in developing culturally-aware MLLMs. Our work thus underscores an important research direction for developing more culturally adaptive and responsible MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。