构建毒性机制推理基准,评估模型是否真懂毒理路径。
ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway
- 基于毒性通路(AOP)设计多器官毒性推理任务
- 模型预测准确但解释常违背生物机制
- 引入机制感知训练可提升推理与预测性能
大语言模型在分子性质预测中表现突出,但毒性源于复杂生物机制,仅靠化学结构难以可靠预测。现有基准未能系统评估模型的机制推理能力,导致模型生成流畅却不具生物学真实性的解释。为此,我们提出ToxReason,一个基于有害结局通路(Adverse Outcome Pathway, AOP)的基准,用于评估跨多器官的器官水平毒性推理能力。该基准融合实验药物-靶点相互作用证据与毒性标签,要求模型从分子起始事件(MIE)推断至最终有害结局(AO)。通过ToxReason,我们评估了多种LLMs的毒性预测性能与推理质量,发现强预测能力并不等同于可靠推理。进一步表明,机制感知训练能有效提升机制推理能力,并改善毒性预测表现。结果强调,在毒性建模中必须将机制推理融入评估与训练流程以实现可信预测。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have enabled molecular reasoning for property prediction. However, toxicity arises from complex biological mechanisms beyond chemical structure, necessitating mechanistic reasoning for reliable prediction. Despite its importance, current benchmarks fail to systematically evaluate this capability. LLMs can generate fluent but biologically unfaithful explanations, making it difficult to assess whether predicted toxicities are grounded invalid mechanisms. To bridge this gap, we introduce ToxReason, a benchmark grounded in the Adverse Outcome Pathway (AOP) that evaluates organ-level toxicity reasoning across multiple organs. ToxReason integrates experimental drug-target interaction evidence with toxicity labels, requiring models to infer both toxic outcomes and their underlying mechanisms from Molecular Initiating Event (MIE) to Adverse Outcome (AO). Using ToxReason, we evaluate toxicity prediction performance and reasoning quality across diverse LLMs. We find that strong predictive performance does not necessarily imply reliable reasoning. Furthermore, we show that reasoning-aware training improves mechanistic reasoning and, consequently, toxicity prediction performance. Together, these results underscore the necessity of integrating reasoning into both evaluation and training for trustworthy toxicity modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。