arXiv:2603.03407cs.CL2026-03中稿 · , Learning Meaning…

揭秘大模型如何存储药物分类知识

Tracing Pharmacological Knowledge In Large Language Models

  • 用激活修补和线性探测分析模型内部药物语义表征
  • 早期层和药物名中间词元是关键信息载体,非末尾词元
  • 药物语义分布于多个词元,非单个词元存储

大型语言模型(LLMs)在药物发现任务中表现出色,但其内部如何编码药理学知识仍不明确。本研究通过因果分析与探测方法,探究基于Llama的生物医学语言模型中药物类别语义的表示与检索机制。采用激活修补定位药物类别信息在模型各层与词元位置的存储情况,并结合基于词元级与求和池化激活的线性探测。结果表明,早期层对药物类别知识编码起关键作用,最强因果效应来自药物类别跨度内的中间词元,而非末尾词元。线性探测进一步显示,药理学语义分布于多个词元,在嵌入空间中已存在,词元级探测接近随机水平,而求和池化表示达到最高准确率。综合表明,药物类别语义并非集中于单一词元,而是由分布式表示构成。该研究首次系统揭示了大模型中药理学知识的内在机制,为理解生物医学语义在大模型中的编码方式提供新视角。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown strong empirical performance across pharmacology and drug discovery tasks, yet the internal mechanisms by which they encode pharmacological knowledge remain poorly understood. In this work, we investigate how drug-group semantics are represented and retrieved within Llama-based biomedical language models using causal and probing-based interpretability methods. We apply activation patching to localize where drug-group information is stored across model layers and token positions, and complement this analysis with linear probes trained on token-level and sum-pooled activations. Our results demonstrate that early layers play a key role in encoding drug-group knowledge, with the strongest causal effects arising from intermediate tokens within the drug-group span rather than the final drug-group token. Linear probing further reveals that pharmacological semantics are distributed across tokens and are already present in the embedding space, with token-level probes performing near chance while sum-pooled representations achieve maximal accuracy. Together, these findings suggest that drug-group semantics in LLMs are not localized to single tokens but instead arise from distributed representations. This study provides the first systematic mechanistic analysis of pharmacological knowledge in LLMs, offering insights into how biomedical semantics are encoded in large language models.

大模型药理学可解释性分布式表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。