用语言模型的不确定性生成跨语言表示,更准且无缺失值。
Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations
- 以单语模型预测熵作为语言结构相似性指标
- 在典型分类和多语言任务中表现接近现有方法
- 适合需要动态、密集语言表征的研究者
我们提出Entropy2Vec,一种通过利用单语语言模型的熵来构建跨语言表示的新框架。与传统类型学库存存在特征稀疏和静态快照的问题不同,Entropy2Vec利用语言模型中的固有不确定性来捕捉语言间的类型学关系。通过在单一语言上训练语言模型,我们假设其预测熵反映了该语言与其他语言的结构相似性:熵越低,相似度越高;熵越高,差异越大。该方法生成稠密、不稀疏的语言嵌入,可适应不同时间尺度且无缺失值。实证评估表明,Entropy2Vec嵌入与已知类型学类别一致,并在下游多语言自然语言处理任务中取得与LinguAlchemy框架相当的表现。
原文摘要 · Abstract (English)
We introduce Entropy2Vec, a novel framework for deriving cross-lingual language representations by leveraging the entropy of monolingual language models. Unlike traditional typological inventories that suffer from feature sparsity and static snapshots, Entropy2Vec uses the inherent uncertainty in language models to capture typological relationships between languages. By training a language model on a single language, we hypothesize that the entropy of its predictions reflects its structural similarity to other languages: Low entropy indicates high similarity, while high entropy suggests greater divergence. This approach yields dense, non-sparse language embeddings that are adaptable to different timeframes and free from missing values. Empirical evaluations demonstrate that Entropy2Vec embeddings align with established typological categories and achieved competitive performance in downstream multilingual NLP tasks, such as those addressed by the LinguAlchemy framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。