latent knowledge能预测大模型学新知识的速度和泛化能力
Latent Knowledge as a Predictor of Fact Acquisition in Fine-Tuned Large Language Models
- 用随机解码检测隐含知识,发现其是快速学习的关键预测因子
- 微调后HPO准确率从2.8%升至71.9%,但未见知识泛化达5.8%
- 有隐含知识的模型更不易遗忘,尤其对训练中强化过的事实
大型语言模型在预训练后存储生物医学事实的能力不均:部分事实存在于权重中但难以通过确定性解码获取(即隐含知识),另一些则几乎未被表示。我们对Llama 3.1 8B Instruct模型进行微调,学习人类表型本体(HPO,800对)和基因本体(GO,400对训练数据)的标识符映射,预留400个GO对用于测试泛化能力。将学习视为20个训练轮次中的时间事件过程,采用随机解码检测基线隐含知识,并使用Cox比例风险模型识别知识获取、泛化与退化的预测因子。基线时HPO确定性召回率为2.8%,微调后提升至71.9%。隐含知识是更快获取事实的最强预测因子(风险比2.6),并与更高、更早的学习峰值速率及更快收敛相关;标识符频率和注释数量影响较小。对未见的GO事实泛化罕见(5.8%),但若存在隐含知识则可能性更高。已知的GO映射在未见术语上退化更频繁,表明训练期间的强化具有保护作用。结果表明,隐含知识可预测微调期间事实学习速度及对未见本体事实的有限泛化能力,而抗退化则取决于事实是否在训练中被强化。
原文摘要 · Abstract (English)
Large language models store biomedical facts with uneven strength after pretraining: some facts are present in the weights but are not reliably accessible under deterministic decoding (latent knowledge), while others are scarcely represented. We fine tuned Llama 3.1 8B Instruct to learn ontology term identifier mappings from the Human Phenotype Ontology (800 pairs) and the Gene Ontology (400 training pairs), withholding 400 GO pairs to test generalization. Treating learning as a time to event process across 20 epochs, we used stochastic decoding to detect latent knowledge at baseline and Cox proportional hazards models to identify predictors of acquisition, generalization, and degradation. Baseline deterministic recall for HPO was 2.8%, rising to 71.9% after fine-tuning. Latent knowledge was the strongest predictor of faster fact acquisition (HR 2.6) and was associated with earlier, higher peak learning rates and faster convergence; identifier frequency and curated annotation counts had smaller effects. Generalization to withheld GO facts was uncommon (5.8%) but more likely when latent knowledge was present. Previously correct GO mappings degraded more often for withheld (unseen) terms than for trained (seen) terms, suggesting a protective effect of reinforcement during training. These results show that latent knowledge predicts both the speed of factual learning during fine-tuning and the limited generalization of unseen ontology facts, while resistance to degradation depends on whether facts are reinforced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。