用近红外光谱和机器学习区分韩牛与荷斯坦牛肉,防伪造假。
Harnessing Near-Infrared Spectroscopy and Machine Learning for Traceable Classification of Hanwoo and Holstein Beef
- 通过便携式近红外光谱仪采集700-1100nm波段数据,非破坏性检测。
- 随机森林模型准确率最高,AUC达0.8826,神经网络召回率最优为0.7804。
- 适合食品溯源、防伪检测场景,尤其对高敏感度需求的质检系统有参考价值。
本研究评估了近红外光谱(NIRS)结合先进机器学习(ML)技术在区分韩牛牛肉(HNB)与荷斯坦牛肉(HLB)中的应用,以应对食品真实性、标签错误及掺假问题。采用便携式NIRS设备,在700至1100 nm波长范围内获取快速、无损的吸收光谱数据。从本地超市采集共40个腰大肌样本,均等分配于两类牛肉。主成分分析(PCA)显示两种牛肉具有显著不同的光谱特征,可解释总方差的93.72%。对比包括LDA、SVM、LR、随机森林、梯度提升(GB)、KNN、决策树(DT)、朴素贝叶斯(NB)和神经网络(NN)在内的多种机器学习模型,经超参数调优与5折交叉验证后,随机森林表现最佳,受试者工作特征曲线下面积(ROC AUC)达0.8826,紧随其后的是SVM模型(AUC 0.8747)。GB与NN模型表现良好,交叉验证得分分别为0.752。其中,神经网络模型达到最高召回率0.7804,表明其在高灵敏度场景中适用性更强。而决策树与朴素贝叶斯表现相对较弱。逻辑回归与支持向量机在准确率、精确率与召回率之间取得较好平衡,成为优选方案。结果表明,将NIRS与机器学习结合可为肉类真实性提供强大可靠的检测方法,显著助力食品欺诈识别。
原文摘要 · Abstract (English)
This study evaluates the use of Near-Infrared spectroscopy (NIRS) combined with advanced machine learning (ML) techniques to differentiate Hanwoo beef (HNB) and Holstein beef (HLB) to address food authenticity, mislabeling, and adulteration. Rapid and non-invasive spectral data were attained by a portable NIRS, recording absorbance data within the wavelength range of 700 to 1100 nm. A total of 40 Longissimus lumborum samples, evenly split between HNB and HLB, were obtained from a local hypermarket. Data analysis using Principal Component Analysis (PCA) demonstrated distinct spectral patterns associated with chemical changes, clearly separating the two beef varieties and accounting for 93.72% of the total variance. ML models, including Linear Discriminant Analysis (LDA), Support Vector Machine (SVM), Logistic Regression (LR), Random Forest, Gradient Boosting (GB), K-Nearest Neighbors, Decision Tree (DT), Naive Bayes (NB), and Neural Networks (NN), were implemented, optimized through hyperparameter tuning, and validated by 5-fold cross-validation techniques to enhance model robustness and prevent overfitting. Random Forest provided the highest predictive accuracy with a Receiver Operating Characteristic (ROC) Area Under the Curve (AUC) of 0.8826, closely followed by the SVM model at 0.8747. Furthermore, GB and NN algorithms exhibited satisfactory performances, with cross-validation scores of 0.752. Notably, the NN model achieved the highest recall rate of 0.7804, highlighting its suitability in scenarios requiring heightened sensitivity. DT and NB exhibited comparatively lower predictive performance. The LR and SVM models emerged as optimal choices by effectively balancing high accuracy, precision, and recall. This study confirms that integrating NIRS with ML techniques offers a powerful and reliable method for meat authenticity, significantly contributing to detecting food fraud.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。