arXiv:2604.14815cs.CL2026-04

用芬兰病理报告微调BERT,预测预训练收益。

Domain Fine-Tuning FinBERT on Finnish Histopathological Reports: Train-Time Signals and Downstream Correlations

论文配图:Domain Fine-Tuning FinBERT on Finnish Histopathological Reports: Train-Time Signals and Downstream Correlations
图 1 · 摘自论文原文
  • 在无标签芬兰医学文本上微调FinBERT,观察嵌入变化
  • 通过嵌入几何变化预测领域预训练的潜在收益
  • 适合医疗AI数据稀缺场景下的模型优化

在标注数据极少的NLP分类任务中,对变压器模型进行无标签数据上的领域微调是成熟方法。本文有两个目标:(1) 描述我们在芬兰医学文本上微调芬兰BERT模型的观察结果;(2) 探索通过观察领域微调导致的嵌入几何变化,来预测芬兰BERT领域特定预训练的潜在收益。我们的主要动机是医疗AI中的普遍困境——获取数据集,尤其是标注数据,常常面临长时间延迟。

原文摘要 · Abstract (English)

In NLP classification tasks where little labeled data exists, domain fine-tuning of transformer models on unlabeled data is an established approach. In this paper we have two aims. (1) We describe our observations from fine-tuning the Finnish BERT model on Finnish medical text data. (2) We report on our attempts to predict the benefit of domain-specific pre-training of Finnish BERT from observing the geometry of embedding changes due to domain fine-tuning. Our driving motivation is the common\situation in healthcare AI where we might experience long delays in acquiring datasets, especially with respect to labels.

医疗AI领域微调嵌入分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。