揭示大模型中多语言表示的可分离性与层级结构,助理解跨语言干扰机制。
When Language Representations Interact: Separability and Cross-Lingual Effects in LLMs

- 用因果几何方法分析28组双语对比,发现语言在调整协方差后具线性可分性
- 同语系语言(如日耳曼、罗曼语族)呈现单纯形几何结构,反映层级关系
- 结果可指导可信部署,避免监控或干预时引发意外跨语言效应
大型语言模型具备强大的多语言能力,但其内部表示难以解释。理解这些交互对保障多语言系统的可靠行为至关重要。已有研究显示,因果-几何结构可解释某些概念在近似线性且可分离方向上的编码方式,但该框架是否适用于语言身份相关且具有层级结构的多语言模型仍不明确。本研究对三种多语言大模型中的28组双语对比进行因果-几何分析,探究语言何时表现为近似独立因素,何时仍存在结构化依赖。结果表明,语言概念在协方差调整后的因果内积下具有稳定的线性表示,且可分性良好;结构偏差反映了语言相似性。同一语系(如日耳曼语族或罗曼语族)的语言呈现单纯形几何结构,暗示其层级组织。这些发现将因果-几何可解释性拓展至多语言场景,揭示了可分离性与相似性在多语言表示中的共存机制,为预测概念间结构依赖提供了依据。这对可信部署具有重要意义:语言间的残余结构可能导致模型监控或干预时出现非预期的跨语言影响。
原文摘要 · Abstract (English)
Large language models exhibit strong multilingual capabilities, however, their internal representations are difficult to interpret. Understanding these interactions is important for ensuring reliable behavior in multilingual systems. Recent work has shown that causal-geometric structure can explain how certain concepts are encoded as approximately linear and separable directions, but whether this framework extends to multilingual models, where language identity is correlated and hierarchical, is underexplored. We apply causal-geometric analysis to multilingual LLMs, studying 28 bilingual contrasts across three models, allowing us to analyze when languages behave as approximately independent factors and when structured dependencies persist. We find evidence that language concepts admit stable linear representations that are largely separable under a covariance-adjusted (causal) inner product, with structured deviations reflecting linguistic similarity. Moreover, languages within the same family (such as Germanic or Romance) exhibit a simplex-like geometric structure, suggesting hierarchical organization. These results extend causal-geometric interpretability to multilingual settings and provide insight into how separability and similarity may exist in multilingual LLM representations, motivating interpretability analyses that diagnose when and how structured dependencies between concepts can be anticipated. This has implications for trustworthy deployment, as residual structure between languages may lead to unintended cross-lingual effects when models are monitored or intervened upon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。