用可解释的符号回归分析肠道菌群数据,发现能同时兼顾预测效果与生物意义。
Interpreting Microbiome Relative Abundance Data Using Symbolic Regression
- 采用符号回归提取菌群相对丰度的显式数学关系。
- 在超1万样本上达到与随机森林相当的准确率(F1分数未明确给出)。
- 结果可帮助理解复杂模型,适合临床与生物机制研究者使用。
解析微生物组内部复杂互作对开发诊断和治疗策略至关重要。传统机器学习模型常缺乏可解释性,难以提供临床与生物学洞见。本文探索符号回归(SR)在微生物组相对丰度数据中的应用,聚焦结直肠癌(CRC)。SR以高可解释性著称,与随机森林、梯度提升决策树等模型对比,基于F1分数和准确率等指标评估。研究整合71项研究,涵盖超过10,000个样本及749个物种特征。结果显示,SR在预测性能上表现合理,且在模型可解释性方面显著占优,可生成显式数学表达式,揭示微生物组内的潜在生物关系,为临床与生物学分析提供关键支持。实验还表明,SR可用于知识蒸馏,辅助理解XGBoost等复杂模型。代码已开源,便于复现与后续研究。
原文摘要 · Abstract (English)
Understanding the complex interactions within the microbiome is crucial for developing effective diagnostic and therapeutic strategies. Traditional machine learning models often lack interpretability, which is essential for clinical and biological insights. This paper explores the application of symbolic regression (SR) to microbiome relative abundance data, with a focus on colorectal cancer (CRC). SR, known for its high interpretability, is compared against traditional machine learning models, e.g., random forest, gradient boosting decision trees. These models are evaluated based on performance metrics such as F1 score and accuracy. We utilize 71 studies encompassing, from various cohorts, over 10,000 samples across 749 species features. Our results indicate that SR not only competes reasonably well in terms of predictive performance, but also excels in model interpretability. SR provides explicit mathematical expressions that offer insights into the biological relationships within the microbiome, a crucial advantage for clinical and biological interpretation. Our experiments also show that SR can help understand complex models like XGBoost via knowledge distillation. To aid in reproducibility and further research, we have made the code openly available at https://github.com/swag2198/microbiome-symbolic-regression .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。