构建首个英法学术文献跨语言检索数据集,提升非英语科研内容可发现性。
CLIRudit: Cross-Lingual Information Retrieval of Scientific Documents
- 基于多语言元数据,用英文关键词匹配法语摘要
- 不依赖翻译的稠密向量检索效果接近翻译系统
- 适合关注跨语言学术搜索与开放获取的研究者
跨语言信息检索(CLIR)帮助用户在非母语文本中查找所需文献,尤其在学术搜索中意义重大。本文提出CLIRudit,一个基于加拿大出版平台Érudit构建的英法学术检索数据集。利用多语言元数据,将英文作者撰写的关键词作为查询,对应非英文摘要作为目标文档,该方法可推广至其他语言和资源库。我们在多种稀疏与稠密检索模型上进行基准测试,对比有无机器翻译的情况。结果表明:无需翻译的稠密嵌入检索性能接近使用机器翻译的系统;文档翻译比查询翻译更有效;稀疏检索器在文档翻译后仍具竞争力且效率更高。除发布首个英法学术检索数据集外,还提供可复现的基准测试方法,以提升非英语学术内容的可及性。
原文摘要 · Abstract (English)
Cross-lingual information retrieval (CLIR) helps users find documents in languages different from their queries. This is especially important in academic search, where key research is often published in non-English languages. We present CLIRudit, a novel English-French academic retrieval dataset built from Érudit, a Canadian publishing platform. Using multilingual metadata, we pair English author-written keywords as queries with non-English abstracts as target documents, a method that can be applied to other languages and repositories. We benchmark various first-stage sparse and dense retrievers, with and without machine translation. We find that dense embeddings without translation perform nearly as well as systems using machine translation, that translating documents is generally more effective than translating queries, and that sparse retrievers with document translation remain competitive while offering greater efficiency. Along with releasing the first English-French academic retrieval dataset, we provide a reproducible benchmarking method to improve access to non-English scholarly content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。