用强化学习优化无标签单细胞数据的基因标志物筛选。
Knowledge-Guided Biomarker Identification for Label-Free Single-Cell RNA-Seq Data: A Reinforcement Learning Perspective
- 融合多算法知识构建初始搜索边界,指导基因筛选方向。
- 通过专家行为设计奖励函数,动态优化基因面板选择。
- 相比传统方法更精准高效,适合生物标志物发现场景。
基因面板选择旨在从无标签基因组数据中识别最具信息量的基因生物标志物。传统方法依赖领域知识、嵌入式机器学习模型或基于启发式的迭代优化,常引入偏差和低效问题,可能掩盖关键生物学信号。为此,我们提出一种迭代基因面板选择策略:首先利用现有基因选择算法的集成知识建立初步边界或先验信息,引导初始搜索空间;随后引入强化学习,通过专家行为设计的奖励函数实现基因面板的动态精炼与靶向选择。该整合方法缓解了初始边界带来的偏差,同时发挥强化学习的随机适应性优势。大量对比实验、案例研究及下游分析表明,该方法在无标签生物标志物发现中具有更高精度与效率,凸显其在单细胞基因组数据分析中的应用潜力。
原文摘要 · Abstract (English)
Gene panel selection aims to identify the most informative genomic biomarkers in label-free genomic datasets. Traditional approaches, which rely on domain expertise, embedded machine learning models, or heuristic-based iterative optimization, often introduce biases and inefficiencies, potentially obscuring critical biological signals. To address these challenges, we present an iterative gene panel selection strategy that harnesses ensemble knowledge from existing gene selection algorithms to establish preliminary boundaries or prior knowledge, which guide the initial search space. Subsequently, we incorporate reinforcement learning through a reward function shaped by expert behavior, enabling dynamic refinement and targeted selection of gene panels. This integration mitigates biases stemming from initial boundaries while capitalizing on RL's stochastic adaptability. Comprehensive comparative experiments, case studies, and downstream analyses demonstrate the effectiveness of our method, highlighting its improved precision and efficiency for label-free biomarker discovery. Our results underscore the potential of this approach to advance single-cell genomics data analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。