arXiv:2608.20440cs.LGcs.AI2026-08

用拉曼光谱和物理引导的AI,仅靠少量特征就能精准识别食用油真伪。

Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach

  • 结合决策树与非负最小二乘法,从1866维光谱中提取关键特征。
  • 纯油样本分类准确率100%,仅需4个特征,占原始信息0.21%。
  • 适合便携式设备和边缘计算,支持低功耗食品质量监测。

食用油在加工食品中的真实性认证对食品安全、防欺诈及合规至关重要。本研究构建了拉曼光谱与机器学习融合框架,实现光谱内在结构关联、可解释分类与物理引导人工智能(PI-AI)的集成。针对五种食用油,分别在纯态及薯片基质中进行分析,采用t-SNE、K-means聚类、决策树与基于非负最小二乘(NNLS)的光谱分解。无监督分析显示,纯油样品具有更强的类别组织性与可分性,而食品基质导致显著光谱重叠。决策树在纯油样本上仅使用原始1866维光谱中的4个变量即达100%分类准确率;该四变量组合在预剪枝与后剪枝模型中均一致出现,仅占约0.21%的光谱信息,但保持完整测试性能。对于含基质样本,基于NNLS的PI-AI光谱分解有效分离出油信号与纸张、马铃薯贡献。优化后的后剪枝模型在去除纸张与同时去除纸张及马铃薯的样本上,准确率分别达86.4%与85.4%,重要变量减少至5个和4个。四特征表示将数据量减少99.44%且不损失精度。结果表明,通过物理意义明确、高度紧凑且可解释的光谱表征,可实现高精度拉曼油品识别,为轻量化AI、边缘智能、便携传感与嵌入式食品质量监控提供坚实基础。

原文摘要 · Abstract (English)

Authentication of edible oils in processed foods is important for food quality, fraud prevention, and regulatory compliance. This study establishes an integrated Raman spectroscopy and machine-learning framework that links intrinsic spectral organization, interpretable classification, and Physics-Informed Artificial Intelligence (PI-AI). Five edible oils were investigated in pure form and within a fried-potato-chip matrix using t-SNE, K-means clustering, Decision Trees, and Non-Negative Least Squares (NNLS)-based spectral decomposition. Unsupervised analyses revealed substantially stronger class organization and separability in pure oils, whereas food-matrix effects introduced pronounced spectral overlap. Decision Trees achieved 100% classification accuracy for pure oils using only four Raman variables from the original 1866-feature spectral space. These four variables, consistently identified by both pre-pruned and post-pruned models, represented only approximately 0.21% of the available spectral information while retaining perfect test-set performance. For matrix-containing samples, NNLS-based PI-AI spectral decomposition substantially improved classification by separating oil-related signatures from paper and potato contributions. Optimized post-pruned models achieved accuracies of 86.4% and 85.4% for paper-subtracted and paper-plus-potato-subtracted datasets, respectively, while reducing the number of important Raman variables to only five and four. The compact four-feature representation further reduced the data footprint by 99.44% without loss of classification accuracy. Collectively, these findings demonstrate that accurate Raman-based oil identification can be achieved through physically meaningful, highly compact, and interpretable spectral representations, providing a promising foundation for Frugal AI, Edge AI, portable sensing, and embedded food-quality monitoring.

拉曼光谱食用油识别边缘AI可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。