arXiv:2602.23199cs.AI2026-02被引 1

为单细胞语言模型设计统一评估框架,提升生物合理性与可解释性。

SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented Evaluation

  • 构建虚拟细胞抽象统一评估任务,涵盖五类生物学推理需求。
  • 知识增强评估使判断更符合生物学事实,判别力强于传统方法。
  • 适合关注单细胞生物智能、AI可解释性的研究者使用。

大语言模型在科学研究中日益普及,但在单细胞生物学领域,通用与专业模型的评估仍不充分:现有基准零散、任务格式脱离实际应用(如多选题)、评估指标缺乏可解释性与生物学依据。本文提出SC-Arena,一个专为单细胞基础模型设计的自然语言评估框架。该框架引入虚拟细胞抽象,统一表示细胞内在属性与基因级交互。在此范式下,定义了五类自然语言任务:细胞类型注释、图像描述生成、文本生成、扰动预测与科学问答,以检验核心生物学推理能力。针对传统字符串匹配指标的脆弱性,引入知识增强评估,融合外部本体、标记基因数据库与科学文献,实现生物合理且可解释的判断。实验表明:(i) 在虚拟细胞统一评估下,当前模型在需机制或因果理解的复杂任务上表现不均;(ii) 知识增强评估确保生物正确性,提供证据支撑的可解释理由,具备高区分度,克服了传统指标的脆弱与不透明问题。SC-Arena为单细胞生物领域的语言模型评估提供了统一、可解释的框架,推动发展对齐生物学、可泛化的基础模型。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly applied in scientific research, offering new capabilities for knowledge discovery and reasoning. In single-cell biology, however, evaluation practices for both general and specialized LLMs remain inadequate: existing benchmarks are fragmented across tasks, adopt formats such as multiple-choice classification that diverge from real-world usage, and rely on metrics lacking interpretability and biological grounding. We present SC-ARENA, a natural language evaluation framework tailored to single-cell foundation models. SC-ARENA formalizes a virtual cell abstraction that unifies evaluation targets by representing both intrinsic attributes and gene-level interactions. Within this paradigm, we define five natural language tasks (cell type annotation, captioning, generation, perturbation prediction, and scientific QA) that probe core reasoning capabilities in cellular biology. To overcome the limitations of brittle string-matching metrics, we introduce knowledge-augmented evaluation, which incorporates external ontologies, marker databases, and scientific literature to support biologically faithful and interpretable judgments. Experiments and analysis across both general-purpose and domain-specialized LLMs demonstrate that (i) under the Virtual Cell unified evaluation paradigm, current models achieve uneven performance on biologically complex tasks, particularly those demanding mechanistic or causal understanding; and (ii) our knowledge-augmented evaluation framework ensures biological correctness, provides interpretable, evidence-grounded rationales, and achieves high discriminative capacity, overcoming the brittleness and opacity of conventional metrics. SC-Arena thus provides a unified and interpretable framework for assessing LLMs in single-cell biology, pointing toward the development of biology-aligned, generalizable foundation models.

单细胞评估框架知识增强大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。