分析多模型文本图像表示的语义信息分布与预测关系。
A quantitative analysis of semantic information in deep representations of text and images
- 用信息不平衡度量衡量表示间的预测能力,高效计算高维空间互信息。
- 语义信息在中间层最集中,且英语表示比其他语言更具预测性。
- 跨模态预测最强的层因模型类型而异,支持语义收敛假设。
近期观察发现,处理相同或语义相关输入的不同模型表示趋于对齐。本文采用信息不平衡(Information Imbalance)——一种基于秩的非对称度量——量化一个表示预测另一个的能力,作为交叉熵的高效代理,在高维空间中可快速计算。通过测量 DeepSeek-V3 在六种语言对翻译任务中生成表示的信息不平衡,发现语义信息分散于多个词元,且在中间层预测能力最强,具有鲁棒性。我们观测到显著的信息不对称:英语表示系统性更易被其他语言表示预测;且 DeepSeek-V3 表示比小模型 Llama3-8b 更具预测性。在视觉领域,自回归模型的语义信息集中于中间层,编码器模型则集中在最后层,这些层也与图像描述的文本表示呈现最强的跨模态可预测性。结果支持语言、模态和架构间存在语义收敛的假设,同时表明表示间的定向预测能力强烈依赖于层深度、模型规模和语言。
原文摘要 · Abstract (English)
It was recently observed that the representations of different models that process identical or semantically related inputs tend to align. We analyze this phenomenon using the Information Imbalance, an asymmetric rank-based measure that quantifies the capability of a representation to predict another, providing a proxy of the cross-entropy which can be computed efficiently in high-dimensional spaces. By measuring the Information Imbalance between representations generated by DeepSeek-V3 processing translations, we find that semantic information is spread across many tokens, and that semantic predictability is strongest in a set of central layers of the network, robust across six language pairs. We measure clear information asymmetries: English representations are systematically more predictive than those of other languages, and DeepSeek-V3 representations are more predictive of those in a smaller model such as Llama3-8b than the opposite. In the visual domain, we observe that semantic information concentrates in middle layers for autoregressive models and in final layers for encoder models, and these same layers yield the strongest cross-modal predictability with textual representations of image captions. Our results support the hypothesis of semantic convergence across languages, modalities, and architectures, while showing that directed predictability between representations varies strongly with layer-depth, model scale, and language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。