arXiv:2606.05616cs.CL2026-06

大模型靠药名前缀推断药效,可能误判虚构药物。

What's in a Name? Morphological Shortcuts by LLMs in Pharmacology

论文配图:What's in a Name? Morphological Shortcuts by LLMs in Pharmacology
图 1 · 摘自论文原文
  • 用真实词根构造假药名,测试模型是否仅凭前缀判断药性。
  • 653种药分析显示,模型90%以上依赖前缀推断,易混淆相似前缀药物。
  • 早期到中期层的激活模式揭示了这种形貌捷径机制。

单词的形态形式常暗示其意义,但过度依赖此类映射在高风险领域可能导致错误。例如,在医学领域,大模型可仅凭药名前缀(如 wugcillin)自信推断虚构药物的药理特性,并生成看似合理的临床内容。本文通过行为与机制研究,揭示大模型在药理学中的“词缀启发法”现象。利用由真实词根构成的虚构药名,我们发现仅词缀信号即可引发模型对药物类别的整体药理反应。为此提出一个框架,用于判断模型对药物语义的依赖主要来自词缀、词干或整个名称。在653种药物上的应用表明,模型大多通过词缀线索推断药效,却很少明确体现此依赖,甚至会错误关联共享词缀的药物属性。跨模型的激活修补分析进一步将该行为定位至早期到中期网络层。这些发现表明,形态捷径虽隐蔽但可测量,对医疗安全构成实质风险。

原文摘要 · Abstract (English)

The morphological form of a word can often give cues to its meaning, but purely relying on these mappings can lead to overgeneralization in high-stakes domains. In the medical domain, for instance, LLMs can confidently reason about fictitious drugs from their affixes alone (e.g., wugcillin) and generate plausible-looking clinical content. We present a behavioral and mechanistic study of LLM "affix heuristics" in pharmacology. Using fictitious drug names built from real affixes, we show that affix signals alone elicit class-level pharmacological responses. We introduce a framework for identifying whether a model's drug semantics are driven mainly by the affix, the stem, or the drug name as a whole. Applied across 653 drugs, our framework reveals that models often induce drug meaning primarily through affix cues, yet rarely explicitly indicate this reliance, and sometimes incorrectly conflate properties among affix-sharing drugs. Activation patching across models further localizes this behavior to early-mid layers. These findings show that morphological shortcuts pose a subtle but measurable risk to safety.

大模型药理推理形貌捷径安全风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。