用强化学习选基因,兼顾预测准确与生物通路意义。
Biological Pathway Guided Gene Selection Through Collaborative Reinforcement Learning
- 用多智能体强化学习,让基因协作优化预测与通路相关性。
- 在多个数据集上,预测准确率和生物可解释性均显著提升。
- 适合关注基因筛选可解释性的生物医学研究者使用。
高维基因组数据中的基因选择对理解疾病机制和改善治疗结果至关重要。传统特征选择方法虽能识别预测性基因,但常忽略复杂的生物通路与调控网络,导致签名不稳定且缺乏生物学意义。现有方法如基于Lasso的方法或统计过滤,要么只关注单个基因与结局的关联,要么无法捕捉通路层面的交互作用。为解决这一问题,我们提出一种两阶段框架,通过多智能体强化学习(MARL)融合统计选择与生物通路知识。第一阶段采用基于KEGG通路信息的路径引导预筛选策略,实现初步降维。第二阶段将基因建模为协同智能体,在MARL框架中优化预测能力与生物相关性。框架通过图神经网络构建状态表示,设计结合预测性能、基因中心性与通路覆盖率的奖励机制,并利用共享记忆和集中式评判器实现协同学习。在多个基因表达数据集上的实验表明,该方法显著提升了预测准确率与生物可解释性。
原文摘要 · Abstract (English)
Gene selection in high-dimensional genomic data is essential for understanding disease mechanisms and improving therapeutic outcomes. Traditional feature selection methods effectively identify predictive genes but often ignore complex biological pathways and regulatory networks, leading to unstable and biologically irrelevant signatures. Prior approaches, such as Lasso-based methods and statistical filtering, either focus solely on individual gene-outcome associations or fail to capture pathway-level interactions, presenting a key challenge: how to integrate biological pathway knowledge while maintaining statistical rigor in gene selection? To address this gap, we propose a novel two-stage framework that integrates statistical selection with biological pathway knowledge using multi-agent reinforcement learning (MARL). First, we introduce a pathway-guided pre-filtering strategy that leverages multiple statistical methods alongside KEGG pathway information for initial dimensionality reduction. Next, for refined selection, we model genes as collaborative agents in a MARL framework, where each agent optimizes both predictive power and biological relevance. Our framework incorporates pathway knowledge through Graph Neural Network-based state representations, a reward mechanism combining prediction performance with gene centrality and pathway coverage, and collaborative learning strategies using shared memory and a centralized critic component. Extensive experiments on multiple gene expression datasets demonstrate that our approach significantly improves both prediction accuracy and biological interpretability compared to traditional methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。