arXiv:2510.19217cs.CL2025-10Conference of the …被引 2

为跨语言迁移设计匹配语言类型的精准距离度量

Modality Matching Matters: Calibrating Language Distances for Cross-Lingual Transfer in URIEL+

  • 按语言类型定制表示:地理用加权分布,谱系用双曲嵌入,类型学用隐变量模型
  • 在零样本迁移任务中,相关类型距离提升显著,综合距离在多数任务上表现更优
  • 适合需要跨语言知识迁移的研究者,尤其关注语言结构差异的场景

现有语言知识库如 URIEL+ 为跨语言迁移提供了宝贵的地理、谱系和类型学距离,但存在两大局限:其一,统一向量表示难以适应语言数据的多样化结构;其二,缺乏将这些信号整合为单一综合评分的合理方法。本文提出一种类型匹配的语言距离框架,为每种距离类型设计结构感知表示:地理采用说话人加权分布,谱系使用双曲嵌入,类型学采用隐变量模型。我们将这些信号统一为一个稳健、任务无关的复合距离。在多个零样本迁移基准测试中,当距离类型与任务相关时,我们的表示显著提升迁移性能;而复合距离在大多数任务中均带来性能增益。

原文摘要 · Abstract (English)

Existing linguistic knowledge bases such as URIEL+ provide valuable geographic, genetic and typological distances for cross-lingual transfer but suffer from two key limitations. First, their one-size-fits-all vector representations are ill-suited to the diverse structures of linguistic data. Second, they lack a principled method for aggregating these signals into a single, comprehensive score. In this paper, we address these gaps by introducing a framework for type-matched language distances. We propose novel, structure-aware representations for each distance type: speaker-weighted distributions for geography, hyperbolic embeddings for genealogy, and a latent variables model for typology. We unify these signals into a robust, task-agnostic composite distance. Across multiple zero-shot transfer benchmarks, we demonstrate that our representations significantly improve transfer performance when the distance type is relevant to the task, while our composite distance yields gains in most tasks.

跨语言迁移语言距离知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。