arXiv:2604.11490cs.AIcs.CL2026-04

让视觉语言模型更懂不同地区文化,同时不丢全球通用能力。

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

论文配图:Anthropogenic Regional Adaptation in Multimodal Vision-Language Model
图 1 · 摘自论文原文
  • 提出区域适应新范式,兼顾本地化与全局性能。
  • 在东南亚测试中文化相关性提升5%-15%,全球性能保持98%以上。
  • 方法简单有效,适合需本土化部署的AI系统开发者。

尽管视觉语言(VL)领域在跨语言、跨域的图文信息融合方面取得显著进展,但尚缺乏专门评估人类中心对齐的框架。本文提出两个贡献:首先,引入「人为区域适应」(Anthropogenic Regional Adaptation)范式,旨在优化模型在特定区域语境下的相关性,同时保留全局泛化能力;其次,提出名为「地理泛化简易化」(GG-EZ)的简单而有效的适配方法,通过区域数据筛选与模型合并实现。在三种视觉语言架构上——大型视觉语言模型、文生图扩散模型、视觉语言嵌入模型——以及东南亚(SEA)地区的案例研究中,实验表明该范式至关重要,且GG-EZ方法表现优异:在东南亚文化相关性指标上提升5%-15%,同时维持超过98%的全球性能,甚至偶尔超越。研究确立了人为区域对齐作为多模态视觉语言模型跨区域应用的基础范式,并提供了一种兼具区域适配与全局泛化的实用基准方法。

原文摘要 · Abstract (English)

While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and domains, there is still no dedicated framework for assessing human-centric alignment in vision-language systems. We offer two contributions to address this gap. First, we introduce Anthropogenic Regional Adaptation: a novel paradigm that aims to optimize model relevance to specific regional contexts while ensuring the retention of global generalization capabilities. Second, we present a simple, but effective adaptation method named Geographical-generalization-made-easy (GG-EZ), which utilizes regional data filtering and model merging. Through comprehensive experiments on 3 VL architectures: large vision-language models, text-to-image diffusion models, and vision-language embedding models, and a case study in Southeast Asia (SEA) regional adaptation, we demonstrate the importance of Anthropogenic Regional Adaptation and the effectiveness of GG-EZ, showing 5-15% gains in cultural relevance metrics across SEA while maintaining over 98% of global performance and even occasionally surpassing it. Our findings establish Anthropogenic Regional Alignment as a foundational paradigm towards applicability of multimodal vision-language models in diverse regions and demonstrate a simple-yet-effective baseline method that optimizes regional value alignment while preserving global generalization.

视觉语言区域适应文化对齐模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。