大模型医学翻译跨语言能力验证,低资源语言表现不输高资源
Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages
- 用五层框架评估4个大模型在8种语言上的医学翻译
- 704组翻译中语义保留率超92%,高低资源语言差异不显著
- 多模型互验结果一致,适合医疗AI落地与语言公平研究
语言障碍影响美国2730万非英语使用者,但专业医疗翻译成本高且难获取。我们评估了四个前沿大模型(GPT-5.1、Claude Opus 4.5、Gemini 3 Pro、Kimi K2)将22份医学文档翻译成8种语言的表现,涵盖高资源(西班牙语、中文、俄语、越南语)、中资源(韩语、阿拉伯语)和低资源(他加禄语、海地克里奥尔语)。使用五层验证框架,在704组翻译对中,所有模型均实现高语义保留(LaBSE > 0.92),高低资源语言间无显著差异(p = 0.066)。跨模型回译验证排除了同模型循环性干扰(delta = -0.0009)。四模型间一致性高(LaBSE: 0.946),低资源语言中英文术语保留与忠实度无相关性(rho = +0.018, p = 0.82)。结果表明,前沿大模型在不同资源语言下均能有效保持医学含义,对改善医疗语言可及性具有重要启示。
原文摘要 · Abstract (English)
Language barriers affect 27.3 million U.S. residents with non-English language preference, yet professional medical translation remains costly and often unavailable. We evaluated four frontier large language models (GPT-5.1, Claude Opus 4.5, Gemini 3 Pro, Kimi K2) translating 22 medical documents into 8 languages spanning high-resource (Spanish, Chinese, Russian, Vietnamese), medium-resource (Korean, Arabic), and low-resource (Tagalog, Haitian Creole) categories using a five-layer validation framework. Across 704 translation pairs, all models achieved high semantic preservation (LaBSE greater than 0.92), with no significant difference between high- and low-resource languages (p = 0.066). Cross-model back-translation confirmed results were not driven by same-model circularity (delta = -0.0009). Inter-model concordance across four independently trained models was high (LaBSE: 0.946), and lexical borrowing analysis showed no correlation between English term retention and fidelity scores in low-resource languages (rho = +0.018, p = 0.82). These converging results suggest frontier LLMs preserve medical meaning across resource levels, with implications for language access in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。