arXiv:2412.11758cs.IR2024-12被引 2

为东帝汶语构建首个信息检索基础,包含停用词、词干提取器和测试集。

Establishing a Foundation for Tetun Ad-Hoc Text Retrieval: Stemming, Indexing, Retrieval, and Ranking

  • 开发东帝汶语专用停用词表与词干提取器,支持文本预处理
  • 去除连字符和撇号后仅检索标题,性能提升31.37%(效率)
  • 公开发布含59个主题的测试集,助力后续研究

在互联网与数字平台中高效检索信息需要有效的检索方案,但东帝汶语尚无此类系统,导致难以找到相关文档。为填补这一空白,本研究聚焦于东帝汶语的即兴检索任务,构建了基础语言资源:包括停用词列表、词干提取器及测试集。通过评估文档标题与内容的不同策略发现,去除连字符和撇号后仅检索标题,相比基线表现更优:效率提升31.37%,在MAP@10上相对增益+9.40%,在NDCG@10上达+30.35%(使用DFR BM25)。超出前10位的排名点,Hiemstra LM在多种策略与指标下均表现强劲。本工作贡献包括Labadain-Stopwords(160个停用词)、Labadain-Stemmer(三种变体)以及Labadain-Avaliadór测试集(59个主题,33,550份文档,5,900个相关性标注)。所有资源均已公开,供未来研究使用。

原文摘要 · Abstract (English)

Searching for information on the internet and digital platforms requires effective retrieval solutions. However, such solutions are not yet available for Tetun, making it difficult to find relevant documents for search queries in this language. To address this gap, we investigate Tetun text retrieval with a focus on the ad-hoc retrieval task. The study begins with the development of essential language resources -- including a list of stopwords, a stemmer, and a test collection -- that serve as a foundation for Tetun text retrieval. Various strategies are evaluated using document titles and content. The results show that retrieving document titles, after removing hyphens and apostrophes but without applying stemming, improves performance compared to the baseline. Efficiency increases by 31.37%, while effectiveness achieves an average relative gains of +9.40% in MAP@10 and +30.35% in NDCG@10 with DFR BM25. Beyond the top-10 cutoff point, Hiemstra LM demonstrates strong performance across multiple retrieval strategies and evaluation metrics. The contributions of this work include the development of Labadain-Stopwords (a list of 160 Tetun stopwords), Labadain-Stemmer (a Tetun stemmer with three variants), and Labadain-Avaliadór (a Tetun test collection comprising 59 topics, 33,550 documents, and 5,900 qrels). These resources are publicly available to support future research in Tetun information retrieval.

信息检索东帝汶语文本处理测试集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。