arXiv:2511.21211cs.LG2025-11

用快速特征筛选提升饮食限制相关基因识别的准确性和效率。

Robust gene prioritization for Dietary Restriction via Fast-mRMR Feature Selection techniques

  • 采用Fast-mRMR方法筛选关键非冗余基因特征,简化模型。
  • 在饮食限制基因预测中显著优于现有方法,提升性能。
  • 适合高维组学数据下的基因功能研究,提升可解释性。

基因优先排序(识别可能与生物过程相关的基因)正越来越多地借助人工智能技术解决。然而,现有方法在处理高维且标注不全的生物医学数据时表现不佳。本文提出一种更鲁棒、高效的分析流程,利用Fast-mRMR特征选择技术保留相关且非冗余的特征,构建更简单、可解释性更强、计算更高效的分类模型。在饮食限制(Dietary Restriction, DR)这一关注领域中,实验表明该方法显著优于现有方法,并支持整合异构生物特征集以提升性能,而此前此类策略常因噪声累积导致效果下降。研究聚焦于DR,因其具备经验证的注释数据和专家知识用于验证;但该流程可推广至其他生物过程,证明特征选择对高维组学数据中可靠基因优先排序至关重要。

原文摘要 · Abstract (English)

Gene prioritization (identifying genes potentially associated with a biological process) is increasingly tackled with Artificial Intelligence. However, existing methods struggle with the high dimensionality and incomplete labelling of biomedical data. This work proposes a more robust and efficient pipeline that leverages Fast-mRMR Feature Selection to retain only relevant, non-redundant features for classifiers, building simpler, more interpretable and more efficient models. Experiments in our domain of interest, prioritizing genes related to Dietary Restriction (DR), show significant improvements over existing methods and enables us to integrate heterogeneous biological feature sets for better performance, a strategy that previously degraded performance due to noise accumulation. This work focuses on DR given the availability of curated data and expert knowledge for validation, yet this pipeline would be applicable to other biological processes, proving that feature selection is critical for reliable gene prioritization in high-dimensional omics.

基因优先排序特征选择组学分析饮食限制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。