arXiv:2603.11244q-bio.GNcs.LG2026-03

构建标准化评估框架,让基因表达生成模型可比可复现

A Standardized Framework For Evaluating Gene Expression Generative Models

  • 提供分布度量与生物导向分析的统一计算空间
  • 不同实现下指标差异显著,凸显标准化必要性
  • 适合基因组学研究者与生成模型开发者使用

单细胞基因表达数据生成模型的快速发展迫切需要标准化评估框架。当前评估实践存在度量实现不一致、超参数选择不可比、缺乏生物学依据度量等问题。我们提出开源Python框架GGE,提供全面的分布度量,支持明确的计算空间选项,并通过差异表达基因(DEG)分析和扰动效应相关性实现生物动机评估,支持标准化报告与可复现基准测试。通过对单细胞生成建模文献的广泛分析,我们发现目前尚无标准化评估协议。各方法报告的度量在不同空间中以不同超参数计算,导致结果不可比。我们证明了度量值随实现方式变化显著,凸显标准化的紧迫性。GGE使生成方法间的公平比较成为可能,推动扰动响应预测、细胞身份建模和反事实推断的研究进展。

原文摘要 · Abstract (English)

The rapid development of generative models for single-cell gene expression data has created an urgent need for standardised evaluation frameworks. Current evaluation practices suffer from inconsistent metric implementations, incomparable hyperparameter choices, and a lack of biologically-grounded metrics. We present Generated Genetic Expression Evaluator (GGE), an open-source Python framework that addresses these challenges by providing a comprehensive suite of distributional metrics with explicit computation space options and biologically-motivated evaluation through differentially expressed gene (DEG)-focused analysis and perturbation-effect correlation, enabling standardized reporting and reproducible benchmarking. Through extensive analysis of the single-cell generative modeling literature, we identify that no standardized evaluation protocol exists. Methods report incomparable metrics computed in different spaces with different hyperparameters. We demonstrate that metric values vary substantially depending on implementation choices, highlighting the critical need for standardization. GGE enables fair comparison across generative approaches and accelerates progress in perturbation response prediction, cellular identity modeling, and counterfactual inference.

生成模型基因表达评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。