arXiv:2511.21770q-bio.QMcs.LG2025-11

一站式生物数据平台,自动完成统计与机器学习分析。

Automated Statistical and Machine Learning Platform for Biological Research

  • 融合经典统计与随机森林,自动处理数据并优化模型。
  • 在多个化学数据集上实现高分类准确率,支持特征重要性分析。
  • 无需编程基础的科研人员也能高效使用,适合生物研究者。

生物研究日益依赖计算方法分析实验数据并预测分子性质。现有方法通常需组合多种工具进行统计分析与机器学习,导致工作流低效。我们提出一个集成平台,结合经典统计方法与随机森林分类,实现全面的数据分析,适用于生物科学领域。该平台支持自动超参数优化、特征重要性分析,以及t检验、方差分析(ANOVA)和皮尔逊相关分析等统计测试。通过自动数据预处理、类别编码与基于数据特征的自适应模型配置,弥合了传统统计软件、现代机器学习框架与生物学之间的鸿沟,为无编程经验的研究人员提供统一界面。初步测试在具有不同特征分布的多种化学数据集上评估分类精度。结果表明,将统计严谨性与机器学习可解释性结合,可加速生物发现流程且保持方法学可靠性。平台模块化设计支持未来扩展至更多机器学习算法与统计方法,适用于生物信息学场景。

原文摘要 · Abstract (English)

Research increasingly relies on computational methods to analyze experimental data and predict molecular properties. Current approaches often require researchers to use a variety of tools for statistical analysis and machine learning, creating workflow inefficiencies. We present an integrated platform that combines classical statistical methods with Random Forest classification for comprehensive data analysis that can be used in the biological sciences. The platform implements automated hyperparameter optimization, feature importance analysis, and a suite of statistical tests including t tests, ANOVA, and Pearson correlation analysis. Our methodology addresses the gap between traditional statistical software, modern machine learning frameworks and biology, by providing a unified interface accessible to researchers without extensive programming experience. The system achieves this through automatic data preprocessing, categorical encoding, and adaptive model configuration based on dataset characteristics. Initial testing protocols are designed to evaluate classification accuracy across diverse chemical datasets with varying feature distributions. This work demonstrates that integrating statistical rigor with machine learning interpretability can accelerate biological discovery workflows while maintaining methodological soundness. The platform's modular architecture enables future extensions to additional machine learning algorithms and statistical procedures relevant to bioinformatics.

生物信息机器学习自动化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。