arXiv:2505.08411cs.IR2025-05中稿 · the Short Paper tr…被引 3

让搜索模型学会识别拼音化查询,解决非拉丁语系用户的输入难题。

Lost in Transliteration: Bridging the Script Gap in Neural IR

  • 用混合母语与拉丁化文本微调模型,提升跨文字检索能力。
  • 微调后模型在拉丁化查询上性能接近母语查询,显著缩小脚本差距。
  • 适合关注多语言搜索、低资源语言检索的研究者与工程师。

多数人类语言使用非拉丁字母的书写系统。这些语言的用户为方便输入,常将查询转写为拉丁字母形式(如希腊语的Greeklish、阿拉伯语的Arabizi)。本文发现,当前主流检索系统(包括BGE-M3等多语言稠密嵌入模型)在面对这类转写查询时性能急剧下降,形成“脚本差距”。我们探索将“翻译训练”范式应用于转写场景,通过在母语与拉丁化文本混合数据上进一步微调检索模型,显著提升其跨脚本匹配能力。实验表明,优化后的模型在拉丁化查询上的表现可接近母语查询水平。跨域评估与定性分析显示,转写可能丢失原始查询语义细节,提示需进一步研究。

原文摘要 · Abstract (English)

Most human languages use scripts other than the Latin alphabet. Search users in these languages often formulate their information needs in a transliterated -- usually Latinized -- form for ease of typing. For example, Greek speakers might use Greeklish, and Arabic speakers might use Arabizi. This paper shows that current search systems, including those that use multilingual dense embeddings such as BGE-M3, do not generalise to this setting, and their performance rapidly deteriorates when exposed to transliterated queries. This creates a ``script gap" between the performance of the same queries when written in their native or transliterated form. We explore whether adapting the popular ``translate-train" paradigm to transliterations can enhance the robustness of multilingual Information Retrieval (IR) methods and bridge the gap between native and transliterated scripts. By exploring various combinations of non-Latin and Latinized query text for training, we investigate whether we can enhance the capacity of existing neural retrieval techniques and enable them to apply to this important setting. We show that by further fine-tuning IR models on an even mixture of native and Latinized text, they can perform this cross-script matching at nearly the same performance as when the query was formulated in the native script. Out-of-domain evaluation and further qualitative analysis show that transliterations can also cause queries to lose some of their nuances, motivating further research in this direction.

信息检索多语言跨脚本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。