构建细胞状态条件下的药物作用机制推理评测基准,检验模型是否真懂原理。
PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects
- 基于多细胞背景的基因与化学扰动数据,动态结合知识图谱生成机制解释
- 现有模型预测准确但常靠错误逻辑得正确答案,忽略细胞特异性背景
- 提出新模型PertReasonLM,让答案与机制在特定细胞状态下保持一致
科学领域机器学习评估需区分正确预测与正确原因,尤其在真实分布偏移下。我们提出PertReason,一套用于细胞状态条件下的扰动效应机制推理的基准与框架。核心是PertReasonQA,测试模型在复杂分布变化(如新细胞、未见扰动)下生成机制忠实解释的能力。该基准整合多个细胞背景下的单细胞遗传与化学扰动数据,结合知识图谱,动态依据细胞基线状态调整通路,防止通用记忆。对主流模型的评估显示,其预测准确率与机制推理能力存在系统性差距:模型常以错误逻辑得出正确答案,忽视细胞上下文,或产生方向不一致的机制。作为基准参考,我们提出PertReasonLM——一种大语言模型,通过将推理依据锚定于特定细胞通路,强化结果与机制的一致性,以应对上述缺陷。本工作提供了一套诊断工具,用于揭示并缓解数据丰富科学系统中的可靠推理失败问题。
原文摘要 · Abstract (English)
Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under realistic distribution shifts. We introduce PertReason, a knowledge-grounded benchmark and framework suite for cell-state--conditioned reasoning about perturbation effects. At its core, PertReasonQA is a benchmark that tests whether models can generate mechanistically faithful explanations while remaining robust to complex shifts, such as new cells and unseen perturbations. PertReasonQA combines single-cell genetic and chemical perturbation data across multiple cellular contexts with knowledge graphs, and dynamically conditions pathways on cell-specific basal states to avoid generic memorization. Evaluations on state-of-the-art models reveal systematic gaps between predictive accuracy and mechanistic reasoning. Specifically, these models exhibit failure modes largely invisible to standard benchmarks, such as deriving correct answers through flawed logic, ignoring cellular context, and generating directionally inconsistent mechanisms. As a reference probe of the benchmark, we present PertReasonLM, a large language model trained to align outcome predictions with context-specific mechanistic reasoning. Our model targets the identified failure modes by grounding rationales in context-specific pathways and tightening agreement between outcomes and mechanisms. Together, we provide a diagnostic framework for exposing and mitigating failures in faithful reasoning in data-rich scientific systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。