arXiv:2503.22508cs.IR2025-03中稿 · IR4GOOD track in E…被引 1

用零样本语言相似性迁移提升低资源语言检索效果

Improving Low-Resource Retrieval Effectiveness using Zero-Shot Linguistic Similarity Transfer

  • 通过微调神经排序器学习语言变体间的语义相似性
  • 在训练过的变体上提升检索准确率,泛化至未见变体对
  • 适合关注低资源语言信息获取与跨语言检索的研究者

全球化和殖民历史导致世界多数人仅使用英语、法语等少数语言交流,致使奥克语、西西里语等许多语言濒危。这些语言常与高资源语言(如法语、意大利语)共享语法和词汇特征,可聚类为不同程度互懂的语言变体群。当前搜索系统多未针对这些低资源变体训练,用户被迫用高资源语言表达需求,而多数内容也以高资源语言呈现,加剧了低资源语言的检索困境。我们发现现有系统在语言变体间表现不鲁棒,严重影响检索有效性。为此,提出在语言变体对上微调神经排序器,使其利用语言间的语义相似性。实验表明,该方法提升了直接训练变体的性能,并使模型具备更好泛化能力,可应对未见变体对。此外,跨语系迁移效果参差,为未来研究提供方向。

原文摘要 · Abstract (English)

Globalisation and colonisation have led the vast majority of the world to use only a fraction of languages, such as English and French, to communicate, excluding many others. This has severely affected the survivability of many now-deemed vulnerable or endangered languages, such as Occitan and Sicilian. These languages often share some characteristics, such as elements of their grammar and lexicon, with other high-resource languages, e.g. French or Italian. They can be clustered into groups of language varieties with various degrees of mutual intelligibility. Current search systems are not usually trained on many of these low-resource varieties, leading search users to express their needs in a high-resource language instead. This problem is further complicated when most information content is expressed in a high-resource language, inhibiting even more retrieval in low-resource languages. We show that current search systems are not robust across language varieties, severely affecting retrieval effectiveness. Therefore, it would be desirable for these systems to leverage the capabilities of neural models to bridge the differences between these varieties. This can allow users to express their needs in their low-resource variety and retrieve the most relevant documents in a high-resource one. To address this, we propose fine-tuning neural rankers on pairs of language varieties, thereby exposing them to their linguistic similarities. We find that this approach improves the performance of the varieties upon which the models were directly trained, thereby regularising these models to generalise and perform better even on unseen language variety pairs. We also explore whether this approach can transfer across language families and observe mixed results that open doors for future research.

低资源语言检索增强跨语言神经排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。