arXiv:2604.11661cs.LGcs.AI2026-04被引 1

用多智能体和机制图提升虚拟细胞的自主科学推理能力

Towards Autonomous Mechanistic Reasoning in Virtual Cells

论文配图:Towards Autonomous Mechanistic Reasoning in Virtual Cells
图 1 · 摘自论文原文
  • 将生物推理建模为可验证的机制动作图,实现系统化推导
  • 在VC-TRACES数据集上训练后,基因表达预测准确率显著提升
  • 适合做生物机制探索与自动化科学发现的研究者

大语言模型虽被视作加速科学发现的潜力路径,但在生物学等开放领域仍受限于缺乏事实依据且可执行的解释。为此,本文提出一种面向虚拟细胞的结构化解释形式,将生物推理表示为机制动作图,支持系统性验证与证伪。基于此,构建VCR-Agent多智能体框架,融合生物学知识检索与验证器过滤机制,实现自主生成与验证机制推理。利用该框架,发布VC-TRACES数据集,包含来自Tahoe-100M图谱的经验证机制解释。实证表明,使用这些解释进行训练能显著提升事实精确性,并为下游基因表达预测提供更有效的监督信号。结果凸显了通过多智能体与严格验证协同实现可靠机制推理的重要性。

原文摘要 · Abstract (English)

Large language models (LLMs) have recently gained significant attention as a promising approach to accelerate scientific discovery. However, their application in open-ended scientific domains such as biology remains limited, primarily due to the lack of factually grounded and actionable explanations. To address this, we introduce a structured explanation formalism for virtual cells that represents biological reasoning as mechanistic action graphs, enabling systematic verification and falsification. Building upon this, we propose VCR-Agent, a multi-agent framework that integrates biologically grounded knowledge retrieval with a verifier-based filtering approach to generate and validate mechanistic reasoning autonomously. Using this framework, we release VC-TRACES dataset, which consists of verified mechanistic explanations derived from the Tahoe-100M atlas. Empirically, we demonstrate that training with these explanations improves factual precision and provides a more effective supervision signal for downstream gene expression prediction. These results underscore the importance of reliable mechanistic reasoning for virtual cells, achieved through the synergy of multi-agent and rigorous verification.

机制推理虚拟细胞多智能体基因预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。