首个中英双语牙科领域评测基准,助力大模型精准理解牙科知识
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
- 构建中英双语牙科问答与语料库,覆盖16个细分领域
- 14款模型测试显示跨任务与语言性能差距显著
- 领域微调可大幅提升知识密集型任务表现,适合医疗AI研发者
近年来,大语言模型(LLMs)和医学大模型(Med-LLMs)在通用医疗基准上表现出色,但在需要深层领域知识的专科领域(如牙科)仍缺乏系统评估。本文提出DentalBench,首个面向牙科领域的综合性中英双语评测基准。其包含两个核心部分:DentalQA,一个涵盖4项任务、16个牙科子领域的中英双语问答数据集,共36,597个问题;以及DentalCorpus,一个包含337.35百万词的高质量大规模语料库,支持监督微调(SFT)和检索增强生成(RAG)。我们评估了14种主流大模型(含专有、开源及医学专用模型),发现不同任务类型与语言间存在显著性能差异。进一步实验表明,以Qwen-2.5-3B为例,领域适应能显著提升模型在知识密集型与术语敏感型任务上的表现,凸显领域专属评测基准对开发可信医疗大模型的重要性。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) and medical LLMs (Med-LLMs) have demonstrated strong performance on general medical benchmarks. However, their capabilities in specialized medical fields, such as dentistry which require deeper domain-specific knowledge, remain underexplored due to the lack of targeted evaluation resources. In this paper, we introduce DentalBench, the first comprehensive bilingual benchmark designed to evaluate and advance LLMs in the dental domain. DentalBench consists of two main components: DentalQA, an English-Chinese question-answering (QA) benchmark with 36,597 questions spanning 4 tasks and 16 dental subfields; and DentalCorpus, a large-scale, high-quality corpus with 337.35 million tokens curated for dental domain adaptation, supporting both supervised fine-tuning (SFT) and retrieval-augmented generation (RAG). We evaluate 14 LLMs, covering proprietary, open-source, and medical-specific models, and reveal significant performance gaps across task types and languages. Further experiments with Qwen-2.5-3B demonstrate that domain adaptation substantially improves model performance, particularly on knowledge-intensive and terminology-focused tasks, and highlight the importance of domain-specific benchmarks for developing trustworthy and effective LLMs tailored to healthcare applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。