arXiv:2505.16661cs.CL2025-05被引 5

针对日本医药领域打造专用大模型并发布三套评估基准。

A Japanese Language Model and Three New Evaluation Benchmarks for Pharmaceutical NLP

  • 在20亿日文药学语料和80亿英文生物医学语料上持续预训练。
  • 模型在术语密集任务上超越开源模型,接近商业模型表现。
  • 新评测集覆盖知识召回、术语归一化和逻辑一致性,适合医药NLP研究者。

我们基于20亿条日文药学语料和80亿条英文生物医学语料,持续预训练出一个面向日本医药领域的专用语言模型。为实现严谨评估,我们引入三个新基准:基于日本药师执照考试的YakugakuQA;测试跨语言同义词与术语归一化的NayoseQA;以及用于评估成对陈述间一致性推理的SogoCheck。实验对比了开源医疗大模型与商业模型(包括GPT-4o)。结果表明,该领域专用模型优于现有开源模型,并在术语密集和知识型任务上达到与商业模型相当的性能。值得注意的是,即使GPT-4o在SogoCheck任务上表现不佳,说明跨句一致性推理仍是未解挑战。本研究提供的评测体系覆盖事实记忆、词汇变体与逻辑一致,为医药NLP提供了更全面的诊断视角。该工作证明了构建实用、安全、低成本的日语领域模型的可行性,并公开了模型、代码与数据集。

原文摘要 · Abstract (English)

We present a Japanese domain-specific language model for the pharmaceutical field, developed through continual pretraining on 2 billion Japanese pharmaceutical tokens and 8 billion English biomedical tokens. To enable rigorous evaluation, we introduce three new benchmarks: YakugakuQA, based on national pharmacist licensing exams; NayoseQA, which tests cross-lingual synonym and terminology normalization; and SogoCheck, a novel task designed to assess consistency reasoning between paired statements. We evaluate our model against both open-source medical LLMs and commercial models, including GPT-4o. Results show that our domain-specific model outperforms existing open models and achieves competitive performance with commercial ones, particularly on terminology-heavy and knowledge-based tasks. Interestingly, even GPT-4o performs poorly on SogoCheck, suggesting that cross-sentence consistency reasoning remains an open challenge. Our benchmark suite offers a broader diagnostic lens for pharmaceutical NLP, covering factual recall, lexical variation, and logical consistency. This work demonstrates the feasibility of building practical, secure, and cost-effective language models for Japanese domain-specific applications, and provides reusable evaluation resources for future research in pharmaceutical and healthcare NLP. Our model, codes, and datasets are released at https://github.com/EQUES-Inc/pharma-LLM-eval.

医药AI日语模型评测基准语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。