首个日语医学大模型评测基准,推动日语医疗AI发展
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models
- 构建8个模型、20个数据集的多任务评测体系
- 日语理解与医学知识越深,模型表现越好
- 适合日语医疗AI研究者和开发者参考
近年来,日本大型语言模型(LLMs)主要聚焦通用领域,针对日语医学领域的进展较少。一个重要障碍是缺乏全面、大规模的评测基准。此外,评估资源也十分有限。为推动该领域发展,我们提出一个新基准,涵盖4类共8个模型,以及5个任务下的20个日语医学数据集。实验结果表明:(1)对日语理解更深入、医学知识更丰富的模型在日语医学任务中表现更优;(2)即使非专为日语医学设计的模型也能表现出意外良好性能;(3)现有模型在某些日语医学任务中仍有巨大提升空间。我们还提供了适配该基准的评测工具及数据集,均已公开于https://huggingface.co/datasets/Coldog2333/JMedBench,以促进后续研究。
原文摘要 · Abstract (English)
Recent developments in Japanese large language models (LLMs) primarily focus on general domains, with fewer advancements in Japanese biomedical LLMs. One obstacle is the absence of a comprehensive, large-scale benchmark for comparison. Furthermore, the resources for evaluating Japanese biomedical LLMs are insufficient. To advance this field, we propose a new benchmark including eight LLMs across four categories and 20 Japanese biomedical datasets across five tasks. Experimental results indicate that: (1) LLMs with a better understanding of Japanese and richer biomedical knowledge achieve better performance in Japanese biomedical tasks, (2) LLMs that are not mainly designed for Japanese biomedical domains can still perform unexpectedly well, and (3) there is still much room for improving the existing LLMs in certain Japanese biomedical tasks. Moreover, we offer insights that could further enhance development in this field. Our evaluation tools tailored to our benchmark as well as the datasets are publicly available in https://huggingface.co/datasets/Coldog2333/JMedBench to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。