LLM翻译的胸部CT报告被专家和模型各自打分,结果大相径庭。
Blinded Radiologist and LLM-Based Evaluation of LLM-Generated Japanese Translations of Chest CT Reports: Comparative Study
- 用LLM生成日文报告,对比人工编辑版本。
- 专家与模型对翻译质量评价基本不一致,模型偏爱自己生成的内容。
- 教育用途中不能仅靠AI评分,仍需医生把关。
背景:准确翻译放射科报告对多语言研究、临床沟通和教学至关重要,但基于大模型的评估有效性尚不明确。目的:评估大模型生成的日文胸部CT报告在教学中的适用性,并比较放射科医生与大模型作为评判者的效果。方法:分析了150份来自CT-RATE-JPN验证集的胸部CT报告。每份英文报告分别对应人工编辑版与由DeepSeek-V3.2生成的日文版。两位认证放射科医生及一名住院医师独立进行盲评,从术语准确性、可读性、整体质量、放射科风格真实性四个方面评分。同时,三名大模型裁判(DeepSeek-V3.2、Mistral Large 3、GPT-5)也对同一组进行评分。采用QWK与百分比一致性评估一致性。结果:医生间一致性极低(QWK=0.01~0.06),医生与大模型间几乎无共识(QWK=-0.04~0.15)。医生1认为术语相当(59%),更偏好大模型在可读性(51%)和整体质量(51%);医生2认为可读性相当(75%),更倾向人工版整体质量(40% vs 21%)。三名大模型均强烈偏爱大模型输出(70%-99%),且93%以上认为其更像放射科医生写作风格。结论:大模型生成的翻译常显自然流畅,但两名医生评价差异显著,大模型裁判则强烈偏向自身输出且与医生无共识。教学使用时,仅依赖自动化大模型评估不足,仍需专业医生审核。
原文摘要 · Abstract (English)
Background: Accurate translation of radiology reports is important for multilingual research, clinical communication, and radiology education, but the validity of LLM-based evaluation remains unclear. Objective: To evaluate the educational suitability of LLM-generated Japanese translations of chest CT reports and compare radiologist assessments with LLM-as-a-judge evaluations. Methods: We analyzed 150 chest CT reports from the CT-RATE-JPN validation set. For each English report, a human-edited Japanese translation was compared with an LLM-generated translation by DeepSeek-V3.2. A board-certified radiologist and a radiology resident independently performed blinded pairwise evaluations across 4 criteria: terminology accuracy, readability, overall quality, and radiologist-style authenticity. In parallel, 3 LLM judges (DeepSeek-V3.2, Mistral Large 3, and GPT-5) evaluated the same pairs. Agreement was assessed using QWK and percentage agreement. Results: Agreement between radiologists and LLM judges was near zero (QWK=-0.04 to 0.15). Agreement between the 2 radiologists was also poor (QWK=0.01 to 0.06). Radiologist 1 rated terminology as equivalent in 59% of cases and favored the LLM translation for readability (51%) and overall quality (51%). Radiologist 2 rated readability as equivalent in 75% of cases and favored the human-edited translation for overall quality (40% vs 21%). All 3 LLM judges strongly favored the LLM translation across all criteria (70%-99%) and rated it as more radiologist-like in >93% of cases. Conclusions: LLM-generated translations were often judged natural and fluent, but the 2 radiologists differed substantially. LLM-as-a-judge showed strong preference for LLM output and negligible agreement with radiologists. For educational use of translated radiology reports, automated LLM-based evaluation alone is insufficient; expert radiologist review remains important.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。