arXiv:2507.04432q-bio.MNcs.CL2025-07被引 1

用小模型精准补全结核病相关分子调控路径,准确率超80%。

Reconstructing Biological Pathways by Applying Selective Incremental Learning to (Very) Small Language Models

  • 基于BERT的小模型通过主动学习筛选关键样本进行增量训练。
  • 仅用25%数据量(约130条)即实现80%以上预测准确率。
  • 高置信度错误样本比低置信度正确样本更利于模型优化。

生成式人工智能在多个领域广泛应用,但通用大语言模型常产生不可靠的“幻觉”答案,限制其在医学与生物医学研究中的应用。本文提出:针对特定任务设计小型、领域专用的语言模型,是更合理的选择。我们采用一个仅约1.1亿参数的微型语言模型,基于BERT架构,执行细胞内调控路径中分子相互作用的预测任务,以填补现有知识空白。通过主动学习策略,仅选取最具信息量的样本进行训练,从人工整理的通路数据库中恢复已知调控关系。实验表明,使用少于25%(约130条)的520条潜在调控关系,即可实现超过80%的预测准确率。进一步发现,以信息熵为指标迭代选择新训练样本时,高置信度的错误陈述(低熵)对提升准确率贡献最大;而正确但低置信度的样本几乎无益,甚至可能拖慢学习速度。

原文摘要 · Abstract (English)

The use of generative artificial intelligence (AI) models is becoming ubiquitous in many fields. Though progress continues to be made, general purpose large language AI models (LLM) show a tendency to deliver creative answers, often called "hallucinations", which have slowed their application in the medical and biomedical fields where accuracy is paramount. We propose that the design and use of much smaller, domain and even task-specific LM may be a more rational and appropriate use of this technology in biomedical research. In this work we apply a very small LM by today's standards to the specialized task of predicting regulatory interactions between molecular components to fill gaps in our current understanding of intracellular pathways. Toward this we attempt to correctly posit known pathway-informed interactions recovered from manually curated pathway databases by selecting and using only the most informative examples as part of an active learning scheme. With this example we show that a small (~110 million parameters) LM based on a Bidirectional Encoder Representations from Transformers (BERT) architecture can propose molecular interactions relevant to tuberculosis persistence and transmission with over 80% accuracy using less than 25% of the ~520 regulatory relationships in question. Using information entropy as a metric for the iterative selection of new tuning examples, we also find that increased accuracy is driven by favoring the use of the incorrectly assigned statements with the highest certainty (lowest entropy). In contrast, the concurrent use of correct but least certain examples contributed little and may have even been detrimental to the learning rate.

生物通路小模型主动学习结核病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。