arXiv:2512.17678cs.LGcs.AI2025-12中稿 · Transactions on Ma…

一次训练选出关键基因,让预测和选基因同时优化。

You Only Train Once: Differentiable Subset Selection for Omics Data

  • 端到端可微框架,预测任务直接指导基因选择
  • 在两个数据集上表现优于现有方法,基因子集更紧凑有效
  • 适合需要精简基因标志物的生物标记物发现场景

从单细胞转录组数据中选取紧凑且信息丰富的基因子集,对生物标志物发现、提高可解释性及低成本检测至关重要。然而,现有特征选择方法多为多阶段流程或依赖事后特征归因,导致选择与预测耦合弱。本文提出 YOTO(You Only Train Once)——一种端到端可微框架,联合识别离散基因子集并完成预测。模型中,预测任务直接引导基因选择,而学习到的基因子集反过来塑造预测表示,形成闭环反馈,使模型在训练过程中迭代优化选择内容与预测能力。不同于以往方法,YOTO 强制稀疏性,仅被选基因参与推理,无需额外训练下游分类器。通过多任务学习设计,模型在相关任务间共享表示,使部分标注数据可相互促进,在不增加训练步骤的情况下发现跨任务泛化的基因子集。我们在两个代表性单细胞 RNA-seq 数据集上评估 YOTO,结果表明其持续优于当前最优基线。该方法通过稀疏、端到端、多任务基因子集选择,提升了预测性能,并获得紧凑且有意义的基因子集,推动了生物标志物发现与单细胞分析的发展。

原文摘要 · Abstract (English)

Selecting compact and informative gene subsets from single-cell transcriptomic data is essential for biomarker discovery, improving interpretability, and cost-effective profiling. However, most existing feature selection approaches either operate as multi-stage pipelines or rely on post hoc feature attribution, making selection and prediction weakly coupled. In this work, we present YOTO (you only train once), an end-to-end framework that jointly identifies discrete gene subsets and performs prediction within a single differentiable architecture. In our model, the prediction task directly guides which genes are selected, while the learned subsets, in turn, shape the predictive representation. This closed feedback loop enables the model to iteratively refine both what it selects and how it predicts during training. Unlike existing approaches, YOTO enforces sparsity so that only the selected genes contribute to inference, eliminating the need to train additional downstream classifiers. Through a multi-task learning design, the model learns shared representations across related objectives, allowing partially labeled datasets to inform one another, and discovering gene subsets that generalize across tasks without additional training steps. We evaluate YOTO on two representative single-cell RNA-seq datasets, showing that it consistently outperforms state-of-the-art baselines. These results demonstrate that sparse, end-to-end, multi-task gene subset selection improves predictive performance and yields compact and meaningful gene subsets, advancing biomarker discovery and single-cell analysis.

基因筛选单细胞端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。