arXiv:2602.05692cs.CL2026-02ACL被引 3

首个多语言医学错误检测与修正基准,覆盖中英阿三语临床文本。

MedErrBench: A Fine-Grained Multilingual Benchmark for Medical Error Detection and Correction with Clinical Expert Annotations

  • 基于十类常见错误类型构建跨语言标注数据集
  • 非英语场景下模型表现显著低于英语,暴露语言短板
  • 适合医疗AI安全评估、多语言NLP研究者使用

现有或生成的临床文本中的错误可能引发严重后果,尤其是误诊或错误治疗建议。随着大语言模型在医疗应用中的广泛使用,建立专用评估基准至关重要。然而,此类数据集仍稀缺,尤其在多语言和多样化语境下。本文提出MedErrBench,首个由临床专家指导的多语言医学错误检测、定位与修正基准。涵盖英语、阿拉伯语和中文,基于十类常见错误类型,采用真实临床案例并经领域专家标注与审核。我们评估了多种通用、语言特定及医学领域语言模型在三项任务上的表现。结果揭示显著性能差距,尤其在非英语环境中,凸显需构建以临床为根基、具备语言意识的系统。通过公开MedErrBench及评估协议,旨在推动多语言临床NLP发展,促进全球更安全、更公平的AI医疗应用。数据集见附录,匿名版本可在https://github.com/congboma/MedErrBench获取。

原文摘要 · Abstract (English)

Inaccuracies in existing or generated clinical text may lead to serious adverse consequences, especially if it is a misdiagnosis or incorrect treatment suggestion. With Large Language Models (LLMs) increasingly being used across diverse healthcare applications, comprehensive evaluation through dedicated benchmarks is crucial. However, such datasets remain scarce, especially across diverse languages and contexts. In this paper, we introduce MedErrBench, the first multilingual benchmark for error detection, localization, and correction, developed under the guidance of experienced clinicians. Based on an expanded taxonomy of ten common error types, MedErrBench covers English, Arabic and Chinese, with natural clinical cases annotated and reviewed by domain experts. We assessed the performance of a range of general-purpose, language-specific, and medical-domain language models across all three tasks. Our results reveal notable performance gaps, particularly in non-English settings, highlighting the need for clinically grounded, language-aware systems. By making MedErrBench and our evaluation protocols publicly-available, we aim to advance multilingual clinical NLP to promote safer and more equitable AI-based healthcare globally. The dataset is available in the supplementary material. An anonymized version of the dataset is available at: https://github.com/congboma/MedErrBench.

医学AI多语言NLP错误检测临床文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。