为卢森堡语构建了跨语言句子嵌入模型,提升低资源语言表现
LuxEmbedder: A Cross-Lingual Approach to Enhanced Luxembourgish Sentence Embeddings
- 用人工标注的平行语料训练跨语言嵌入模型
- 在自建的卢森堡语语义相似度任务上表现优于基线
- 证明低资源语言并行数据对其他低资源语言有益
句子嵌入模型在主题建模、文档聚类和推荐系统等自然语言处理任务中至关重要。然而,这些模型严重依赖双语平行数据,而卢森堡语等低资源语言的平行语料稀缺,导致其单语和跨语言嵌入模型性能不佳。为此,我们构建了一个小但高质量的人工生成的跨语言平行语料库,用于训练 LuxEmbedder——一个增强型卢森堡语句子嵌入模型,具备强跨语言能力。此外,我们发现将低资源语言纳入并行训练数据,对其他低资源语言的性能提升效果甚至优于仅使用高资源语言对。鉴于低资源语言缺乏句子嵌入评测基准,我们还创建了首个针对卢森堡语的释义检测基准,旨在填补空白并推动后续研究。
原文摘要 · Abstract (English)
Sentence embedding models play a key role in various Natural Language Processing tasks, such as in Topic Modeling, Document Clustering and Recommendation Systems. However, these models rely heavily on parallel data, which can be scarce for many low-resource languages, including Luxembourgish. This scarcity results in suboptimal performance of monolingual and cross-lingual sentence embedding models for these languages. To address this issue, we compile a relatively small but high-quality human-generated cross-lingual parallel dataset to train LuxEmbedder, an enhanced sentence embedding model for Luxembourgish with strong cross-lingual capabilities. Additionally, we present evidence suggesting that including low-resource languages in parallel training datasets can be more advantageous for other low-resource languages than relying solely on high-resource language pairs. Furthermore, recognizing the lack of sentence embedding benchmarks for low-resource languages, we create a paraphrase detection benchmark specifically for Luxembourgish, aiming to partially fill this gap and promote further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。