arXiv:2511.21075cs.LGcs.AI2025-11被引 4

让大模型更懂生物医学知识,通过识别知识空白提升推理能力

Aligning LLMs with Biomedical Knowledge using Balanced Fine-Tuning

  • 根据生物文本的不确定性特征,设计双尺度微调方法,聚焦知识密集区
  • 在多个生物任务中表现优于传统微调,尤其在稀疏奖励强化学习中持续提升
  • 适合需要精准生物推理与生成的专业研究者,助力药物与基因研究

将大语言模型用于加速生命科学研究,需使其与生物医学知识深度融合。我们发现,生物医学文本的不确定性结构与通用文本截然不同:密集的低置信度片段反映的是认知性知识缺口(如复杂的因果链、罕见实体),而非通用文本中稀疏的随机性风格差异。基于此,提出平衡微调(BFT),一种双尺度后训练方法,结合组归一化标记重加权与序列级样本重分配,优先优化知识密集且具高认知不确定性的样本。在医疗评估、生物推理、稀疏奖励强化学习及生物表征任务中,BFT在相同训练设置下表现更稳定,优于SFT和DFT。当替换GeneAgent(GPT-4o)与VCWorld(Gemini-2.5-Flash)中的默认闭源主干模型时,经BFT对齐的70B模型在生物过程推理与化学扰动预测上表现更优。关键的是,所有BFT变体在后续使用稀疏奖励的GRPO优化后性能进一步提升,而SFT与DFT则下降,表明认知感知的后训练能提供更稳健的策略初始化。此外,BFT对齐模型生成的生物医学简介更准确专业;将其嵌入文本编码模型后,所得表征可支持基因级、细胞级及扰动响应任务,说明其生成能力有助于生物表征学习,并推动更广泛的下游应用。

原文摘要 · Abstract (English)

Engineering LLMs to accelerate life sciences research requires a robust alignment with biomedical knowledge. We observe that biomedical text exhibits a fundamentally different uncertainty structure from general text: dense low-confidence runs encode epistemic knowledge gaps (dense causal chains, rare entities) rather than the sparse aleatoric stylistic variation typical of general text. Based on this discovery, we propose Balanced Fine-Tuning (BFT), a dual-scale post-training method that combines group-normalized token reweighting with sequence-level reallocation toward knowledge-dense samples exhibiting dense epistemic uncertainty. Across medical evaluation, biological reasoning, sparse-reward RL, and biological representation tasks, BFT provides more consistent gains than SFT and DFT under a shared training setup. When replacing the default closed-source backbones in GeneAgent (GPT-4o) and VCWorld (Gemini-2.5-Flash), the BFT-aligned 70B model delivers stronger performance across biological process reasoning and chemical perturbation prediction. Critically, all BFT variants further improve after subsequent GRPO with sparse rewards, while SFT and DFT degrade, suggesting that epistemic-aware post-training provides a more robust policy initialization. Beyond text generation, BFT-aligned LLMs produce more accurate and professional biomedical profile texts; after encoding these profiles with a text embedding model, the resulting representations support gene-level, cell-level, and perturbation-response tasks, suggesting that BFT-enhanced generation can facilitate biological representation and, in turn, broader biomedical downstream tasks.

大模型对齐生物医学认知不确定性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。