揭示CLIP模型训练中纹理-形状偏见的动态演变及其与人类感知和鲁棒性的关系。
On the dynamic evolution of CLIP texture-shape bias and its relationship to human alignment and model robustness
- 按训练阶段追踪CLIP表示特征变化,发现纹理偏好随训练减弱。
- 早期模型更贴近人类低层感知,但对噪声敏感;后期更鲁棒,但感知对齐下降。
- 该现象在不同规模模型中一致,提示存在系统性权衡机制。
对比语言-图像模型如CLIP展现出卓越的泛化能力,但其内部视觉表征在训练过程中的演化以及与人类感知的关系仍不明确。现有分析多聚焦于训练完成后的模型,忽略了表征偏见与感知对齐的动态过程。本文对CLIP模型训练全过程进行逐轮分析,关注纹理-形状偏见、与人类感知判断的对齐度及对图像噪声的敏感性。基于涵盖低层图像质量评估、中层感知相似性、显著性对应和噪声鲁棒性的多个感知基准,我们发现了一致的、依赖训练阶段的表征转变:早期阶段呈现强纹理偏好,与低层人类感知测量高度对齐,且对高斯噪声扰动更敏感;随着训练推进,纹理偏好逐渐减弱,转向更以形状为基础的表征,同时噪声鲁棒性提升,低层感知对齐度下降。这一现象在多个尺寸的CLIP模型中均被观察到,表明其非特定于某一架构规模。研究揭示了感知对齐、特征偏见与鲁棒性在多模态模型训练中的协同演化规律,提出早期低层感知对齐与后期鲁棒性之间存在系统性权衡,为理解视觉-语言模型表征动态及其与人类视觉处理的关系提供了新视角。
原文摘要 · Abstract (English)
Contrastive language-image models such as CLIP have demonstrated remarkable generalization capabilities. However, how their internal visual representations evolve during training and how this evolution relates to human perception remains poorly understood. Most existing analysis characterize fully trained models, leaving the dynamics of representational biases and perceptual alignment largely unexplored. In this work, we present an epoch-by-epoch analysis of CLIP models throughout training, focusing on the evolution of texture-shape bias, alignment with human perceptual judgements, and sensitivity to image noise. Using multiple perceptual benchmarks spanning low-level image quality assessment, mid-level perceptual similarity, saliency correspondence, and noisy robustness, we identify a consistent, training-stage-dependent representational transition. Early training stages exhibit strong texture bias, elevated alignment with low-level human perceptual measures, and increased sensitivity to Gaussian noise perturbations. As training progresses, this texture bias gradually diminishes in favor of more shape-based representations, coinciding with improved robustness to noise and a decline in low-level perceptual alignment. Importantly, these dynamics are consistently observed across multiple CLIP model scales, indicating that the phenomenon is not specific to a particular architecture size. Our findings provide an empirical characterization of how perceptual alignment, feature bias, and robustness co-evolve during multimodal model training. This work reveals a systematic trade-off between early low-level perceptual alignment and later robustness, offering new insights into the representational dynamics of vision-language models and their relationship to human visual processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。