arXiv:2602.05334cs.IR2026-02被引 1

构建中文技术领域跨语言检索评估集,支持中英双语检索测试

NeuCLIRTech: Chinese Monolingual and Cross-Language Information Retrieval Evaluation in a Challenging Domain

论文配图:NeuCLIRTech: Chinese Monolingual and Cross-Language Information Retrieval Evaluation in a Challenging Domain
图 1 · 摘自论文原文
  • 基于2023-2024年TREC NeuCLIR主题构建,含中英双语文档对
  • 110个查询共35,962条相关性判断,具备强区分能力
  • 提供神经排序模型融合基线,适合重排算法开发者使用

为衡量检索技术进展,需具备能真实区分系统性能的相关性标注数据集。本文提出NeuCLIRTech,一个面向技术领域跨语言信息检索的评估集合。该集合包含原始中文技术文档及其机器翻译成英文的版本,涵盖110个查询及35,962条相关性标注,支持中文单语检索与以英文为查询语言的跨语言检索两种场景。NeuCLIRTech整合了TREC NeuCLIR 2023与2024年赛道的主题内容,具备充分的统计判别力。此外,集合还提供了基于强大神经检索系统的融合基线,使重排算法开发者无需依赖BM25作为第一阶段检索器。数据集及相关资源已发布于Hugging Face Datasets。

原文摘要 · Abstract (English)

Measuring advances in retrieval requires test collections with relevance judgments that can faithfully distinguish systems. This paper presents NeuCLIRTech, an evaluation collection for cross-language retrieval over technical information. The collection consists of technical documents written natively in Chinese and those same documents machine translated into English. It includes 110 queries with relevance judgments. The collection supports two retrieval scenarios: monolingual retrieval in Chinese, and cross-language retrieval with English as the query language. NeuCLIRTech combines the TREC NeuCLIR track topics of 2023 and 2024. The 110 queries with 35,962 document judgments provide strong statistical discriminatory power when trying to distinguish retrieval approaches. A fusion baseline of strong neural retrieval systems is included so that developers of reranking algorithms are not reliant on BM25 as their first stage retriever. The dataset and artifacts are released on Huggingface Datasets

信息检索跨语言中文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。