大模型在不同语言间共享语法概念表征,揭示跨语言抽象机制。
Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
- 用稀疏自编码器挖掘模型中跨语言共享的语法特征方向
- 移除多语言特征后分类性能接近随机水平,证明其跨语言性
- 可精准调控机器翻译中的语法行为,适合多语言研究者参考
人类双语者常在相似脑区处理多种语言,取决于第二语言习得时间和熟练度。大语言模型(LLMs)如何学习并编码多种语言?本文研究了诸如语法数、性、时态等形态句法概念在多种语言间的表征共享程度。我们在 Llama-3-8B 和 Aya-23-8B 上训练稀疏自编码器,发现抽象语法概念通常在多个语言间通过共享特征方向进行编码。通过因果干预验证这些表征的多语言特性:仅移除多语言特征即导致跨语言分类性能降至接近随机水平。进一步利用这些特征在机器翻译任务中精确修改模型行为,展示了这些特征在模型中的普遍性与选择性作用。结果表明,即使主要以英语数据训练,模型仍能发展出稳健的跨语言形态句法抽象表征。
原文摘要 · Abstract (English)
Human bilinguals often use similar brain regions to process multiple languages, depending on when they learned their second language and their proficiency. In large language models (LLMs), how are multiple languages learned and encoded? In this work, we explore the extent to which LLMs share representations of morphsyntactic concepts such as grammatical number, gender, and tense across languages. We train sparse autoencoders on Llama-3-8B and Aya-23-8B, and demonstrate that abstract grammatical concepts are often encoded in feature directions shared across many languages. We use causal interventions to verify the multilingual nature of these representations; specifically, we show that ablating only multilingual features decreases classifier performance to near-chance across languages. We then use these features to precisely modify model behavior in a machine translation task; this demonstrates both the generality and selectivity of these feature's roles in the network. Our findings suggest that even models trained predominantly on English data can develop robust, cross-lingual abstractions of morphosyntactic concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。