arXiv:2503.18570cs.IRcs.CL2025-03被引 3
用密集检索技术处理埃塞俄比亚语,突破低资源语言信息检索难题。
Dense Retrieval for Low Resource Languages -- the Case of Amharic Language
- 基于稠密向量构建埃塞俄比亚语检索模型
- 在1200万使用者的语言上实现有效检索
- 为低资源语言研究提供实证参考
本文报告了在埃塞俄比亚语(一种使用人口达1.2亿的低资源语言)上应用密集检索器所遇到的挑战与取得的成果。阿迪斯阿贝巴大学在推动埃塞俄比亚语信息检索过程中所面临的努力与困难将在报告中详细阐述。研究展示了在缺乏标注数据和语言资源的情况下,如何通过稠密检索技术实现有效的跨文档语义匹配,为低资源语言的信息检索提供了可行路径。
原文摘要 · Abstract (English)
This paper reports some difficulties and some results when using dense retrievers on Amharic, one of the low-resource languages spoken by 120 millions populations. The efforts put and difficulties faced by University Addis Ababa toward Amharic Information Retrieval will be developed during the presentation.
低资源语言信息检索稠密检索
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。