arXiv:2507.06210cs.CVcs.CL2025-07被引 6

用合成图像和上下文描述增强CLIP的文化感知能力

CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions

  • 构建合成数据集CulTwin,包含视觉相似但文化不同的概念对
  • 在文化特定任务上提升5.49%细粒度识别准确率,同时保持通用性
  • 适合需要理解跨文化视觉差异的研究与应用

预训练视觉语言模型(如CLIP)在通用多模态理解上表现优异,但难以捕捉依赖上下文的细微视觉线索,导致难以区分外观相似却文化含义不同的概念。这一缺陷主要源于高质量文化数据、上下文信息及负例样本的不足。为此,我们设计了一套数据构建流程,利用开源VLM和文生图模型生成CulTwin——一个包含成对概念-描述-图像三元组的合成文化数据集。这些三元组中的概念视觉相似但文化不同。随后,我们在CulTwin上微调CLIP,开发出CultureCLIP,通过定制对比学习使文化概念与上下文增强的描述及合成图像对齐。在多个文化特定基准测试中,CultureCLIP显著优于基线CLIP,细粒度概念识别最高提升5.49%,同时保持了原始模型的泛化能力,验证了数据合成与模型训练范式的有效性。

原文摘要 · Abstract (English)

Pretrained vision-language models (VLMs) such as CLIP excel in general multimodal comprehension but often struggle to capture nuanced, context-dependent visual cues. This makes it difficult to distinguish between similar-looking concepts with potentially different cultural meanings. Such deficiencies are mainly due to a limited amount of high-quality cultural data, contextual information, and the lack of negative examples that highlight subtle differences. To mitigate this, we design a data curation pipeline leveraging open-sourced VLMs and text-to-image models to construct CulTwin, a synthetic cultural dataset. This dataset consists of paired concept-caption-image triplets, where concepts visually resemble each other but are culturally different. Then, we fine-tune CLIP on CulTwin to develop CultureCLIP, which aligns cultural concepts with contextually enhanced captions and synthetic images through tailored contrastive learning. Experiments on culture-specific benchmarks show that CultureCLIP outperforms the base CLIP, achieving up to a notable 5.49% improvement in fine-grained concept recognition on certain tasks while preserving CLIP's original generalization ability, validating the effectiveness of our data synthesis and VLM backbone training paradigm in capturing subtle cultural distinctions.

文化感知多模态CLIP数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。