arXiv:2412.16894cs.CL2024-12

无监督方法为低资源语言构建双语词典,提升翻译与跨语言任务能力。

Unsupervised Bilingual Lexicon Induction for Low Resource Languages

  • 基于单语嵌入初始化种子词典,通过迭代优化提升双语映射精度。
  • 在英-僧伽罗、英-泰米尔等低资源语言对上验证,显著提升词典准确率。
  • 首次系统评估多种改进技术组合,适合低资源语言研究者参考。

双语词典在自然语言处理中至关重要,但许多低资源语言(LRLs)缺乏此类资源,无法受益于有监督的双语词典构建方法。为此,无监督双语词典构建(UBLI)被提出,其中基于结构的UBLI是主流方法:从单语嵌入中学习初始种子词典,并通过迭代优化逐步改进。尽管已有多种改进策略,但它们常独立实验,未系统对比。本文采用常用结构式框架VecMap的无监督版本,在英-僧伽罗、英-泰米尔、英-旁遮普语三组低资源语言对上开展全面实验,评估不同扩展方法的协同效果,最终确定最优组合。研究还发布了英-僧伽罗和英-旁遮普语的双语词典。

原文摘要 · Abstract (English)

Bilingual lexicons play a crucial role in various Natural Language Processing tasks. However, many low-resource languages (LRLs) do not have such lexicons, and due to the same reason, cannot benefit from the supervised Bilingual Lexicon Induction (BLI) techniques. To address this, unsupervised BLI (UBLI) techniques were introduced. A prominent technique in this line is structure-based UBLI. It is an iterative method, where a seed lexicon, which is initially learned from monolingual embeddings is iteratively improved. There have been numerous improvements to this core idea, however they have been experimented with independently of each other. In this paper, we investigate whether using these techniques simultaneously would lead to equal gains. We use the unsupervised version of VecMap, a commonly used structure-based UBLI framework, and carry out a comprehensive set of experiments using the LRL pairs, English-Sinhala, English-Tamil, and English-Punjabi. These experiments helped us to identify the best combination of the extensions. We also release bilingual dictionaries for English-Sinhala and English-Punjabi.

双语词典低资源语言无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。