构建轻量级多语言伊斯兰文本检索模型,支持真实场景部署。
Efficient and Versatile Model for Multilingual Information Retrieval of Islamic Text: Development and Deployment in Real-World Scenarios
- 融合跨语言与单语训练的混合方法提升检索效果
- 混合模型在多语言场景下表现优异,优于单一训练方式
- 适合需要低成本部署的多语言宗教文本检索应用
尽管多语言信息检索(MLIR)取得进展,但研究与实际部署之间仍存在显著差距。本文利用《古兰经》多语言语料库的独特性,探索为伊斯兰领域构建即席检索系统的方法,以满足多语言用户的信息需求。我们构建了11种检索模型,采用四种训练策略:单语、跨语言、翻译-全量训练,以及一种结合跨语言与单语技术的新混合方法。在域内数据集上的评估表明,混合方法在多种检索场景中表现良好。此外,我们详细分析了不同训练配置对嵌入空间的影响及其对多语言检索有效性的作用。最后,讨论了部署考量,强调使用单一通用轻量模型在真实世界MLIR应用中的成本效益。
原文摘要 · Abstract (English)
Despite recent advancements in Multilingual Information Retrieval (MLIR), a significant gap remains between research and practical deployment. Many studies assess MLIR performance in isolated settings, limiting their applicability to real-world scenarios. In this work, we leverage the unique characteristics of the Quranic multilingual corpus to examine the optimal strategies to develop an ad-hoc IR system for the Islamic domain that is designed to satisfy users' information needs in multiple languages. We prepared eleven retrieval models employing four training approaches: monolingual, cross-lingual, translate-train-all, and a novel mixed method combining cross-lingual and monolingual techniques. Evaluation on an in-domain dataset demonstrates that the mixed approach achieves promising results across diverse retrieval scenarios. Furthermore, we provide a detailed analysis of how different training configurations affect the embedding space and their implications for multilingual retrieval effectiveness. Finally, we discuss deployment considerations, emphasizing the cost-efficiency of deploying a single versatile, lightweight model for real-world MLIR applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。