用大模型权重构建语言度量空间,揭示语言间隐藏联系
Deep Language Geometry: Constructing a Metric Space from LLM Weights
- 通过改进剪枝算法计算权重重要性,自动生成高维语言向量
- 覆盖106种语言,结果与语言家族一致并发现新关联
- 适合语言学、AI交叉研究者,开源代码和向量可直接使用
我们提出一种新框架,利用现代大语言模型(LLMs)内部权重激活构建语言度量空间。不同于依赖人工设计语言特征的传统方法,该方法通过改进的剪枝算法计算权重重要性,自动提取高维向量表示,捕捉反映语言现象的内在特性。我们在多样数据集和多语言LLMs上验证该方法,涵盖106种语言。结果与已知语言家族高度吻合,同时揭示了意外的语言间关联,可能反映历史接触或语言演化关系。源代码、计算出的语言潜在向量及可视化工具已公开,地址为https://github.com/mshamrai/deep-language-geometry。
原文摘要 · Abstract (English)
We introduce a novel framework that utilizes the internal weight activations of modern Large Language Models (LLMs) to construct a metric space of languages. Unlike traditional approaches based on hand-crafted linguistic features, our method automatically derives high-dimensional vector representations by computing weight importance scores via an adapted pruning algorithm. Our approach captures intrinsic language characteristics that reflect linguistic phenomena. We validate our approach across diverse datasets and multilingual LLMs, covering 106 languages. The results align well with established linguistic families while also revealing unexpected inter-language connections that may indicate historical contact or language evolution. The source code, computed language latent vectors, and visualization tool are made publicly available at https://github.com/mshamrai/deep-language-geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。