arXiv:2502.07749q-bio.GNcs.LG2025-02被引 12

机器学习预测细菌表型时,常误判非因果特征,导致结果不可靠。

Whole-Genome Phenotype Prediction with Machine Learning: Open Problems in Bacterial Genomics

  • 用机器学习从全基因组数据预测表型,但易受虚假关联干扰
  • 现有模型高准确率背后隐藏因果推断失效风险
  • 适合关注细菌基因变异与表型因果关系的研究者

如何识别决定细菌性状的因果遗传机制?早期尝试将机器学习用于从基因型预测表型,取得了高准确率。然而,试图从中提取有意义的因果信息时,发现大量被误认为‘因果’的特征实为虚假关联。仅依赖模式识别和相关性在细菌基因组中尤为不可靠,因高维数据和伪相关普遍存在。尽管尚不明确能否突破此障碍,但已有大量研究致力于发现潜在高风险细菌基因变异。本文提出围绕细菌全基因组表型预测及因果效应学习的关键开放问题,并讨论此类数据下机器决策可靠性的挑战。

原文摘要 · Abstract (English)

How can we identify causal genetic mechanisms that govern bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype return high accuracy scores. However, attempts to extract any meaning from the predictive models are found to be corrupted by falsely identified "causal" features. Relying solely on pattern recognition and correlations is unreliable, significantly so in bacterial genomics settings where high-dimensionality and spurious associations are the norm. Though it is not yet clear whether we can overcome this hurdle, significant efforts are being made towards discovering potential high-risk bacterial genetic variants. In view of this, we set up open problems surrounding phenotype prediction from bacterial whole-genome datasets and extending those to learning causal effects, and discuss challenges that impact the reliability of a machine's decision-making when faced with datasets of this nature.

基因组学机器学习因果推断细菌

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。