评测7个开源大模型对日语病理报告的辅助写作能力
Performance Evaluation of Open-Source Large Language Models for Assisting Pathology Report Writing in Japanese
- 从结构化文本生成、错别字修正、医生主观评价三方面评估
- 医学专用模型在逻辑推理和纠错任务中表现更优
- 适合需要辅助写作但无需高度一致解释的临床场景
目前针对日语病理报告撰写的大语言模型性能尚未得到系统评估。本文从三个方面评测了七个开源大语言模型:(A) 按预设格式生成和提取病理诊断文本;(B) 修正日语病理报告中的拼写错误;(C) 由病理科医生和临床医师对模型生成的解释性文本进行主观评价。思维类模型与医学专用模型在需推理的结构化报告任务及错别字修正上表现更优。而解释性文本的偏好在评价者间差异较大。尽管不同任务中模型效用不一,研究结果表明开源大语言模型可在有限但具有临床意义的场景中辅助日语病理报告撰写。
原文摘要 · Abstract (English)
The performance of large language models (LLMs) for supporting pathology report writing in Japanese remains unexplored. We evaluated seven open-source LLMs from three perspectives: (A) generation and information extraction of pathology diagnosis text following predefined formats, (B) correction of typographical errors in Japanese pathology reports, and (C) subjective evaluation of model-generated explanatory text by pathologists and clinicians. Thinking models and medical-specialized models showed advantages in structured reporting tasks that required reasoning and in typo correction. In contrast, preferences for explanatory outputs varied substantially across raters. Although the utility of LLMs differed by task, our findings suggest that open-source LLMs can be useful for assisting Japanese pathology report writing in limited but clinically relevant scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。