arXiv:2501.19264cs.IRcs.CL2025-01中稿 · ECIR 2025被引 12

构建多语言指令跟随检索基准,评估模型跨语言理解能力。

mFollowIR: a Multilingual Benchmark for Instruction Following in Retrieval

  • 基于TREC NeuCLIR叙事数据,设计三语(俄/中/波斯)指令检索任务。
  • 英语训练模型跨语言表现强,但多语言场景性能显著下降。
  • 适合研究多语言信息检索与指令理解的学者参考。

检索系统传统上聚焦于简短且不明确的网页查询。然而,语言模型的进步推动了能够理解复杂多样意图查询的检索模型兴起。目前这些工作仅限于英文,尚不清楚其在其他语言中的表现。本文提出mFollowIR,一个用于衡量检索模型指令遵循能力的多语言基准。该基准基于TREC NeuCLIR叙事(或指令),涵盖俄语、中文和波斯语三种语言,同时提供查询和指令给检索模型。对叙事进行微小改动,以评估模型对细微语义变化的响应能力。报告了多语言(XX-XX)和跨语言(En-XX)两种设置下的性能结果。结果显示,使用指令训练的英语检索器在跨语言任务中表现良好,但在多语言设置下性能明显下降,表明需进一步开发基于指令的多语言检索数据。

原文摘要 · Abstract (English)

Retrieval systems generally focus on web-style queries that are short and underspecified. However, advances in language models have facilitated the nascent rise of retrieval models that can understand more complex queries with diverse intents. However, these efforts have focused exclusively on English; therefore, we do not yet understand how they work across languages. We introduce mFollowIR, a multilingual benchmark for measuring instruction-following ability in retrieval models. mFollowIR builds upon the TREC NeuCLIR narratives (or instructions) that span three diverse languages (Russian, Chinese, Persian) giving both query and instruction to the retrieval models. We make small changes to the narratives and isolate how well retrieval models can follow these nuanced changes. We present results for both multilingual (XX-XX) and cross-lingual (En-XX) performance. We see strong cross-lingual performance with English-based retrievers that trained using instructions, but find a notable drop in performance in the multilingual setting, indicating that more work is needed in developing data for instruction-based multilingual retrievers.

多语言检索指令理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。