首个跨语言医学文本纠错基准,评估大模型在日英语境下的纠错能力。
MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
- 构建日英双语医疗文本纠错三任务基准,自动化生成数据
- 推理型模型纠错效果显著优于普通模型,日语表现差距较小
- 微调后模型超越人类专家,适合医疗AI安全研究者使用
大型语言模型在医疗领域前景广阔,但其在临床文本中检测与纠正错误的能力——确保安全部署的前提——仍缺乏评估,尤其在英语之外的语言中。我们提出MedRECT,一个涵盖日语和英语的跨语言基准,将医疗错误处理分解为错误检测、错误定位(句子提取)和错误修正三个子任务。该基准通过日本医学执照考试题库及人工校对的英文对照,构建了包含663条日语文本(MedRECT-ja)和458条英文文本(MedRECT-en)的数据集,误差与无误样本比例均衡。我们评估了9种主流大模型,涵盖闭源、开源和推理型架构。关键发现:(i) 推理模型显著优于常规架构,错误检测准确率提升最高达13.5%,句子提取提升51.0%;(ii) 跨语言测试显示从英语到日语性能下降5%-10%,但推理模型差异更小;(iii) 针对性LoRA微调使纠错性能提升(日语+0.078,英语+0.168),同时保持推理能力;(iv) 微调后的模型在结构化医疗错误修正任务上超过人类专家表现。据我们所知,MedRECT是首个全面的跨语言医疗错误纠正基准,提供可复现框架与资源,助力多语言医疗大模型的安全开发。
原文摘要 · Abstract (English)
Large language models (LLMs) show increasing promise in medical applications, but their ability to detect and correct errors in clinical texts -- a prerequisite for safe deployment -- remains under-evaluated, particularly beyond English. We introduce MedRECT, a cross-lingual benchmark (Japanese/English) that formulates medical error handling as three subtasks: error detection, error localization (sentence extraction), and error correction. MedRECT is built with a scalable, automated pipeline from the Japanese Medical Licensing Examinations (JMLE) and a curated English counterpart, yielding MedRECT-ja (663 texts) and MedRECT-en (458 texts) with comparable error/no-error balance. We evaluate 9 contemporary LLMs spanning proprietary, open-weight, and reasoning families. Key findings: (i) reasoning models substantially outperform standard architectures, with up to 13.5% relative improvement in error detection and 51.0% in sentence extraction; (ii) cross-lingual evaluation reveals 5-10% performance gaps from English to Japanese, with smaller disparities for reasoning models; (iii) targeted LoRA fine-tuning yields asymmetric improvements in error correction performance (Japanese: +0.078, English: +0.168) while preserving reasoning capabilities; and (iv) our fine-tuned model exceeds human expert performance on structured medical error correction tasks. To our knowledge, MedRECT is the first comprehensive cross-lingual benchmark for medical error correction, providing a reproducible framework and resources for developing safer medical LLMs across languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。