用机器学习预测果胶水解提取参数,准确率超94%。
A Comparative Analysis of Machine Learning Algorithms for Multi-Task Prediction of the Parameters of the Pectin Hydrolysis--Extraction Process

- 采用11种算法对比,CatBoost在多任务回归中表现最佳。
- 原料类型贡献度达63.6%,温度与保温时间次之。
- 成果已转为可部署的交互式网页工具,适合工业研发使用。
本研究针对复杂的多参数果胶水解-提取工艺控制难题,基于1000组实验室实验数据(7类植物原料,4个变量:温度85–130°C、压力0.9–2.2 atm、保温时间3–10 min、pH 1.5–2.0),记录了果胶产率、半乳糖醛酸含量、分子量和酯化度四项输出。比较了11种机器学习算法:正则化线性模型、集成方法(随机森林、梯度提升、XGBoost、CatBoost、Extra Trees)、K近邻、支持向量回归和多层感知机。经超参数优化后,CatBoost平均决定系数R²达0.946,表现最优。特征重要性分析显示,原料类型贡献度占总重要性的63.6%,其次为温度和保温时间。开发的流程已导出为生产就绪格式,并部署为交互式网页界面。结果表明,集成方法结合严谨统计与可解释人工智能,显著减少物理实验需求,为智能果胶生产控制奠定基础。
原文摘要 · Abstract (English)
This study addresses the challenge of controlling a complex, multi-parameter technological process -- pectin hydrolysis--extraction -- using machine learning methods. The experimental foundation is a unique database comprising 1,000 laboratory experiments conducted under controlled conditions on seven types of plant raw material with four variable process factors (temperature 85--130 C, pressure 0.9--2.2 atm, holding time 3--10 min, pH 1.5--2.0). Four output characteristics were recorded: pectin yield, galacturonic acid content, molecular weight, and degree of esterification. To solve the multi-task regression problem, 11 algorithms were trained and compared: regularised linear models, ensemble methods (Random Forest, Gradient Boosting, XGBoost, CatBoost, Extra Trees), k-nearest neighbours, support vector regression, and a multilayer perceptron. The best results were demonstrated by CatBoost (average R-squared approximately 0.946 after hyperparameter optimisation). Feature importance analysis revealed the dominant role of the raw material type (63.6% of total importance), followed by temperature and holding time. The developed pipeline was exported in a production-ready format and deployed as an interactive web interface. The findings demonstrate that ensemble methods combined with rigorous statistical analysis and interpretable AI significantly reduce the need for physical experiments and form the basis for intelligent pectin production control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。