多语言医学信息检索新基准,揭示英日间性能断崖式下降
MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval

- 构建六语言、三任务异构基准,分离跨语言对齐与概念区分能力
- 英语模型在日语上nDCG@10从0.818骤降至0.056,暴露严重跨语言失效
- 适合评估医疗RAG系统的真实多语言泛化能力,尤其关注非英语场景
临床场景中检索增强生成(RAG)日益需要在以英文为主导的证据语料库中进行多语言检索。多语言医学检索需具备三种能力:跨语言对齐、概念区分和证据检索。然而现有基准仅孤立评估这些能力,未衡量生物医学专长与多语言覆盖之间的交互。本文提出MMed-Bench-IR,一个设计用于解耦这些维度的基准,涵盖6种语言和三个结构异质的任务:(1) 基于统一医学语言系统(UMLS)的6,127个跨语言医疗问答检索;(2) 覆盖4,975个混淆集、分三个难度层级的概念区分任务;(3) 针对RAG的2,040个质量保证查询的多语言证据检索。三任务无概念与查询重叠,确保总分反映真实能力广度。对十种系统在六类范式下的评估显示严重跨语言失败:生物医学编码器在英语中nDCG@10达0.818,但在日语中降至0.056,该差距无法被仅含英语的基准检测。
原文摘要 · Abstract (English)
Retrieval-augmented generation (RAG) in clinical settings increasingly requires multilingual retrieval against predominantly English evidence corpora. Multilingual medical retrieval demands three capabilities: cross-lingual alignment, concept discrimination, and evidence retrieval. However, existing benchmarks evaluate these only in isolation, leaving the interaction between biomedical expertise and multilingual coverage unmeasured. We introduce MMed-Bench-IR, a benchmark designed to disentangle these axes across 6 languages and three structurally heterogeneous tasks: (1) cross-lingual medical QA retrieval with 6,127 queries grounded in the Unified Medical Language System (UMLS), (2) concept discrimination over 4,975 confusion sets at three difficulty tiers, and (3) multilingual evidence retrieval for RAG with 2,040 quality-assured queries. The three tasks share zero concept and query overlap by design, ensuring that aggregate scores reflect genuine capability breadth. Evaluation of ten systems across six paradigm families reveals severe cross-lingual failure: biomedical encoders that score 0.818 nDCG@10 in English drop to 0.056 in Japanese, a gap that English-only benchmarks cannot detect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。