用稀疏检索模型SPLARE实现多语言文档搜索,效果超越主流密集模型。
Naver Labs Europe @ WSDM CUP | Multilingual Retrieval
- 基于SPLARE-7B模型,结合轻量级重排序与分数融合提升性能
- 在跨语言检索任务中优于Qwen3-8B-Embed等主流密集模型
- 适合关注多语言信息检索与稀疏模型应用的研究者
本文介绍我们在WSDM Cup 2026多语言文档检索任务中的参与情况。该任务提供了一个具有挑战性的跨语言泛化基准,同时为评估我们近期提出的学习型稀疏检索模型SPLARE提供了自然测试平台。SPLARE能够生成通用的稀疏潜在表示,特别适用于多语言检索场景。我们进行了五次逐步优化的实验,从SPLARE-7B模型出发,引入轻量级改进,包括使用Qwen3-Reranker-4B进行重排序及简单的分数融合策略。结果表明,SPLARE在性能上显著优于如Qwen3-8B-Embed等先进密集基线模型。更广泛地,本提交凸显了学习型稀疏检索模型在非英语主导场景下的持续相关性与竞争力。
原文摘要 · Abstract (English)
This report presents our participation to the WSDM Cup 2026 shared task on multilingual document retrieval from English queries. The task provides a challenging benchmark for cross-lingual generalization. It also provides a natural testbed for evaluating SPLARE, our recently proposed learned sparse retrieval model, which produces generalizable sparse latent representations and is particularly well suited to multilingual retrieval settings. We evaluate five progressively enhanced runs, starting from a SPLARE-7B model and incorporating lightweight improvements, including reranking with Qwen3-Reranker-4B and simple score fusion strategies. Our results demonstrate the strength of SPLARE compared to state-of-the-art dense baselines such as Qwen3-8B-Embed. More broadly, our submission highlights the continued relevance and competitiveness of learned sparse retrieval models beyond English-centric scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。