通过语义对齐实现跨语言模型迁移,用更少数据达到单语模型水平。
Multilinguality as Sense Adaptation
- 基于语义混合与上下文表征的对齐机制,实现跨语言意义传递。
- 在四种语言上性能超越同类方法,目标语言数据仅需1/4仍达单语模型精度。
- 适合低资源语言迁移与多语言模型轻量化部署场景。
本文将多语言能力视为语义适应:通过显式对齐不同语言间的潜在语义表示,而非依赖共享参数和规模。提出SENse-based Symmetric Interlingual Alignment(SENSIA)方法,利用平行语料对背包语言模型进行跨语言适配,同时对目标语言进行语言建模训练以保持流畅性。在四种类型差异显著的语言基准测试中,SENSIA普遍优于现有多语言对齐方法,且在使用2-4倍更少目标语言数据的情况下,达到与从头训练的单语基线相当的准确率。对学习到的语义几何结构分析表明,局部语义拓扑及相对于英语的全局结构基本保留;消融实验显示该方法在设计与规模上均具鲁棒性。
原文摘要 · Abstract (English)
We approach multilinguality as sense adaptation: aligning latent meaning representations across languages rather than relying solely on shared parameters and scale. In this paper, we introduce SENse-based Symmetric Interlingual Alignment (SENSIA), which adapts a Backpack language model from one language to another by explicitly aligning sense-level mixtures and contextual representations on parallel data, while jointly training a target-language language modeling loss to preserve fluency. Across benchmarks on four typologically diverse languages, SENSIA generally outperforms comparable multilingual alignment methods and achieves competitive accuracy against monolingual from-scratch baselines while using 2-4x less target-language data. Analyses of learned sense geometry indicate that local sense topology and global structure relative to English are largely preserved, and ablations show that the method is robust in terms of design and scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。