研究发现医疗模型规模需按任务调整,避免盲目增大。
A Nationwide Japanese Medical Claims Foundation Model: Balancing Model Scaling and Task-Specific Computational Efficiency

- 用5种规模的Transformer在230万患者数据上预训练
- 疾病预测需3200万以上参数,用药预测1100万即饱和
- 比传统模型更准,适合追求效率的医疗AI研发
基于纵向医疗数据的临床风险预测支持个体化医疗。自监督基础模型成为利用大规模未标注医疗记录的有前景方法。在自然语言处理中,缩放定律表明模型越大,预训练损失越低,支持基础模型范式。然而,对于词汇量有限、观测稀疏的结构化医疗数据,增加模型规模是否持续提升下游预测性能尚不明确,因多数研究仅评估单一模型规模。本研究评估了结构化医疗基础模型的模型规模与下游任务性能之间的关系。使用来自日本全国519家医院的230万患者(32家医院随机抽样)的医保数据库,我们对五种规模(220万-10100万参数)的编码器-仅变压器进行疾病发病率和药物预测的预训练。下游性能在任务依赖的阈值处趋于饱和:疾病预测受益于更大模型(3200万-10100万参数),而药物预测在1100万参数时已饱和,使预训练时间减少178小时。所有任务中,表现最佳的模型在精确率-召回率曲线下面积上均优于梯度提升机基线。这些发现表明,与单调下降的预训练损失不同,最优模型规模取决于任务特性。这种任务依赖的饱和现象为平衡结构化医疗基础模型的预测性能与计算成本提供了实用指导。
原文摘要 · Abstract (English)
Clinical risk prediction using longitudinal medical data supports individualized care. Self-supervised foundation models have emerged as a promising approach for leveraging large-scale unlabeled healthcare records. In natural language processing, scaling laws suggest that larger models achieve predictably lower pretraining losses, supporting the foundation model paradigm. However, for structured medical data, characterized by a limited vocabulary and sparse observations, whether increasing model size consistently improves downstream predictions is unclear, as most studies evaluate only a single model scale. In this study, we evaluated the relationship between model scale and downstream task performance for structured medical foundation models. Using a random sample (2.3 million patients, 32 hospitals) from a nationwide 519-hospital Japanese claims database, we pretrained encoder-only Transformers at five scales (2.2M-101M parameters) for disease incidence and medication prediction. Downstream performance saturated at task-dependent thresholds: disease prediction benefited from larger models (32M-101M), whereas medication prediction saturated at 11M, reducing pretraining time by 178 h. Across all tasks, the best-performing model consistently outperformed a Light Gradient Boosting Machine baseline in the area under the precision-recall curve. These findings indicate that, unlike the monotonically decreasing pretraining loss, the optimal model size varied depending on task characteristics. This task-dependent saturation provides practical guidance for balancing predictive performance and computational cost in structured medical foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。