arXiv:2603.13320cs.IRcs.LG2026-03中稿 · and presented at R…

构建尼泊尔语护照服务问答数据集,提升低资源语言信息检索效果

Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications

  • 基于尼泊尔语常见护照问题构建问答数据集
  • 微调SBERT模型在问答匹配中优于传统BM25
  • 多语言E5模型表现最佳,适合低资源语言应用

尼泊尔语作为低资源语言,因缺乏标注数据和计算语言学资源,在构建有效信息检索系统方面面临挑战。本研究通过构建结构化的尼泊尔语问答数据集,聚焦护照相关服务的常见问题,为信息检索模型的训练与评估提供支持。我们对基于Transformer的嵌入模型进行微调,用于问题-答案间的语义相似性匹配,并与基线模型BM25进行比较。此外,还实现了一种融合微调模型与BM25的混合检索方法,并评估其性能。结果表明,微调后的SBERT模型优于BM25,而多语言E5嵌入模型在所有测试模型中表现最佳。

原文摘要 · Abstract (English)

Nepali, a low-resource language, faces significant challenges in building an effective information retrieval system due to the unavailability of annotated data and computational linguistic resources. In this study, we attempt to address this gap by preparing a pair-structured Nepali Question-Answer dataset. We focus on Frequently Asked Questions (FAQs) for passport-related services, building a data set for training and evaluation of IR models. In our study, we have fine-tuned transformer-based embedding models for semantic similarity in question-answer retrieval. The fine-tuned models were compared with the baseline BM25. In addition, we implement a hybrid retrieval approach, integrating fine-tuned models with BM25, and evaluate the performance of the hybrid retrieval. Our results show that the fine-tuned SBERT-based models outperform BM25, whereas multilingual E5 embedding-based models achieve the highest retrieval performance among all evaluated models.

问答系统低资源语言信息检索尼泊尔语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。