对比大模型与传统翻译工具在医学摘要翻译中的表现
Comparing Large Language Models and Traditional Machine Translation Tools for Translating Medical Consultation Summaries: A Pilot Study
- 用标准自动指标评估LLM与传统机器翻译在三种语言的表现
- 传统工具整体更优,尤其复杂文本;大模型在简单中文越南文翻译中表现好
- 提示需领域训练、新评估方法和人工审核,尤其医疗场景
本研究评估大型语言模型(LLMs)与传统机器翻译(MT)工具在将英文医学咨询摘要翻译成阿拉伯语、中文和越南语时的表现。使用患者友好型与临床聚焦型文本,通过标准自动化指标进行评估。结果显示,传统MT工具整体表现更佳,尤其在复杂文本上;而LLMs在简单摘要翻译中表现出潜力,尤其在中文和越南语中。阿拉伯语翻译质量随文本复杂度提升而改善,得益于其形态特征。总体而言,尽管LLMs具备上下文灵活性,但表现仍不稳定,现有评估指标无法捕捉临床相关性。研究强调需开展领域特定训练、改进评估方法,并在医疗翻译中引入人工监督。
原文摘要 · Abstract (English)
This study evaluates how well large language models (LLMs) and traditional machine translation (MT) tools translate medical consultation summaries from English into Arabic, Chinese, and Vietnamese. It assesses both patient, friendly and clinician, focused texts using standard automated metrics. Results showed that traditional MT tools generally performed better, especially for complex texts, while LLMs showed promise, particularly in Vietnamese and Chinese, when translating simpler summaries. Arabic translations improved with complexity due to the language's morphology. Overall, while LLMs offer contextual flexibility, they remain inconsistent, and current evaluation metrics fail to capture clinical relevance. The study highlights the need for domain-specific training, improved evaluation methods, and human oversight in medical translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。