arXiv:2511.06418cs.CL2025-11被引 1

评测大模型对药物机制的理解与推理能力,发现小模型也能媲美大模型。

How Well Do LLMs Understand Drug Mechanisms? A Knowledge + Reasoning Evaluation Dataset

  • 构建新数据集,测试模型在已知机制和反事实情境下的推理能力。
  • o4-mini表现最佳,小模型Qwen3-4B-thinking接近甚至超越其性能。
  • 开放式推理比封闭式更难,中间链路的反事实挑战更大。

药物研发与个性化医疗日益关注预训练大语言模型(LLMs)。在这些领域,模型需具备准确的事实知识及对药物作用机制的深层理解,以在新情境中回忆并推理相关知识。药物作用机制由生物实体间的相互作用构成,形成从药物到靶向疾病的多条因果链。通过组合候选链中的作用效应,可推断药物对特定疾病是否有效。本文提出一个数据集,用于评估模型在已知机制上的知识掌握程度及其在新颖情境下的推理能力,尤其针对训练中罕见的反事实设定。实验表明,o4-mini优于OpenAI的4o、o3及o3-mini模型,而小型模型Qwen3-4B-thinking的表现接近且部分场景超越o4-mini。结果还显示,开放世界推理(需自主召回知识)比封闭世界(提供必要事实)更具挑战性;影响推理链内部环节的反事实任务远难于仅改变药物提示中的外部链接。

原文摘要 · Abstract (English)

Two scientific fields showing increasing interest in pre-trained large language models (LLMs) are drug development / repurposing, and personalized medicine. For both, LLMs have to demonstrate factual knowledge as well as a deep understanding of drug mechanisms, so they can recall and reason about relevant knowledge in novel situations. Drug mechanisms of action are described as a series of interactions between biomedical entities, which interlink into one or more chains directed from the drug to the targeted disease. Composing the effects of the interactions in a candidate chain leads to an inference about whether the drug might be useful or not for that disease. We introduce a dataset that evaluates LLMs on both factual knowledge of known mechanisms, and their ability to reason about them under novel situations, presented as counterfactuals that the models are unlikely to have seen during training. Using this dataset, we show that o4-mini outperforms the 4o, o3, and o3-mini models from OpenAI, and the recent small Qwen3-4B-thinking model closely matches o4-mini's performance, even outperforming it in some cases. We demonstrate that the open world setting for reasoning tasks, which requires the model to recall relevant knowledge, is more challenging than the closed world setting where the needed factual knowledge is provided. We also show that counterfactuals affecting internal links in the reasoning chain present a much harder task than those affecting a link from the drug mentioned in the prompt.

大模型评测药物机制推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。