arXiv:2510.07796cs.LGcs.IR2025-10被引 3

用数学方法提升大模型对药理数据的适应能力

HySim-LLM: Embedding-Weighted Fine-Tuning Bounds and Manifold Denoising for Domain-Adapted LLMs

  • 基于嵌入相似性加权微调,缓解领域偏移问题
  • 理论证明可降低噪声样本对模型性能的影响
  • 适合需要高可靠性的生物医药数据建模场景

从科学文献中提取和标准化药代动力学(PK)信息仍是计算药理学中的重大挑战,限制了数据驱动模型在药物开发中的可靠性。尽管大语言模型(LLMs)在文本理解与推理方面取得显著进展,但其在结构化生物医学数据(如PK表格)上的适应仍受限于异质性、噪声和领域偏移。为此,我们提出HySim-LLM,一个统一的数学与计算框架,整合嵌入加权微调与流形感知去噪,以增强LLMs在该领域的鲁棒性与可解释性。我们建立了两个理论结果:(1) 相似性加权泛化界,量化嵌入偏离下的适应性能;(2) 基于流形的去噪保证,界定噪声或离流形样本的损失贡献。这些定理为结构化生物医学场景下LLM的微调提供了原则性基础。该框架为生物医学及数据密集型科学领域中可靠且可解释的LLM适应提供了数学保障。

原文摘要 · Abstract (English)

The extraction and standardization of pharmacokinetic (PK) information from scientific literature remain significant challenges in computational pharmacology, which limits the reliability of data-driven models in drug development. Large language models (LLMs) have achieved remarkable progress in text understanding and reasoning, yet their adaptation to structured biomedical data, such as PK tables, remains constrained by heterogeneity, noise, and domain shift. To address these limitations, we propose HySim-LLM, a unified mathematical and computational framework that integrates embedding-weighted fine-tuning and manifold-aware denoising to enhance the robustness and interpretability of LLMs. We establish two theoretical results: (1) a similarity-weighted generalization bound that quantifies adaptation performance under embedding divergence, and (2) a manifold-based denoising guarantee that bounds loss contributions from noisy or off-manifold samples. These theorems provide a principled foundation for fine-tuning LLMs in structured biomedical settings. The framework offers a mathematically grounded pathway toward reliable and interpretable LLM adaptation for biomedical and data-intensive scientific domains.

大模型微调生物医学数据去噪理论保障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。