arXiv:2508.09952cs.CLcs.AI2025-08中稿 · ELAMI@MICCAI2025被引 2

针对放射科语言模型,专用分词器提升生成质量并降低计算开销。

Specialised or Generic? Tokenization Choices for Radiology Language Models

  • 对比通用、医学及领域专用分词器在放射科报告摘要任务中的表现。
  • 无预训练时,医学与领域专用分词器显著优于通用分词器;有预训练时差异缩小。
  • 领域专用分词器减少内存占用,适合临床实际部署。

语言模型的词汇表(由分词器定义)对文本生成质量至关重要,但在放射科领域研究不足。本文系统比较了通用、医学及领域专用分词器在三种影像模态上的放射科报告摘要任务表现,并考察了是否在PubMed摘要上进行预训练的影响。结果表明,当模型从零训练时,医学与领域专用词汇表优于通用自然语言分词器;预训练部分缓解了分词器间的性能差异,而领域专用分词器仍表现最佳。此外,领域专用分词器因词汇量小、序列短,显著降低内存需求。这些结果证明,为临床领域定制语言模型词汇表可提升性能并减少计算负担,使其更适用于科研与真实医疗场景。

原文摘要 · Abstract (English)

The vocabulary used by language models (LM) - defined by the tokenizer - plays a key role in text generation quality. However, its impact remains under-explored in radiology. In this work, we address this gap by systematically comparing general, medical, and domain-specific tokenizers on the task of radiology report summarisation across three imaging modalities. We also investigate scenarios with and without LM pre-training on PubMed abstracts. Our findings demonstrate that medical and domain-specific vocabularies outperformed widely used natural language alternatives when models are trained from scratch. Pre-training partially mitigates performance differences between tokenizers, whilst the domain-specific tokenizers achieve the most favourable results. Domain-specific tokenizers also reduce memory requirements due to smaller vocabularies and shorter sequences. These results demonstrate that adapting the vocabulary of LMs to the clinical domain provides practical benefits, including improved performance and reduced computational demands, making such models more accessible and effective for both research and real-world healthcare settings.

语言模型放射科分词器医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。