arXiv:2605.13587stat.MLcs.LG2026-05

用内嵌校准法替代耗时的预处理筛选,显著提速近红外光谱建模。

Reframing preprocessing selection as model-internal calibration in near-infrared spectroscopy: A large-scale benchmark of operator-adaptive PLS and Ridge models

论文配图:Reframing preprocessing selection as model-internal calibration in near-infrared spectroscopy: A large-scale benchmark of operator-adaptive PLS and Ridge models
图 1 · 摘自论文原文
  • 将预处理选择转化为模型内部自适应,避免重复拟合线性模型。
  • 预测精度接近全搜索,但PLS建模时间缩短436至602倍。
  • 适合需要快速建模的工业场景,尤其对实时分析有重要意义。

近红外光谱校准流程中,预处理筛选是最耗时的环节。传统方法通过平滑、导数、去趋势等滤波改变偏最小二乘(PLS)或岭回归所见的光谱方向,但需反复外部搜索并重新拟合几乎相同的线性模型。本文提出将该搜索过程合并为单一校准步骤:对于作用于行谱的严格线性预处理算子A(XA^T),变换后的PLS交叉协方差满足 (XA^T)^T Y = A X^T Y,岭回归依赖于算子诱导的核 X A^T A X^T。这些恒等式使得有限算子库可在模型内部筛选,同时保留原始波长系数;相同机制也适用于廉价评估的线性算子链。针对样本自适应或拟合修正如SNV、MSC、EMSC和ASLS等非严格线性操作,本文证明其边界并作为局部分支保留。研究包含61个回归和17个分类任务,严格配对回归基准为N=32,八种论文变体中,AOM-PLS在简单/最优情况下达到0.991/0.990与0.985/1.002的中位RMSEP比值,分别优于PLS-default/PLS-HPO;AOM-Ridge分别为0.974/0.984和0.918/0.966,优于Ridge-default/Ridge-HPO。AOM-PLS-DA分类器在N=13数据集上中位平衡准确率提升0.159(12/13胜出)。实际效果:PLS-HPO单次运行中位耗时710.81秒,而AOM-PLS仅需1.18–1.63秒,建模时间减少436至602倍。线性算子自适应校准实现了与穷举预处理筛选相当的预测性能,同时使PLS建模时间降低多个数量级。

原文摘要 · Abstract (English)

Preprocessing screening is often the most expensive part of a near-infrared spectroscopy calibration workflow. It works because smoothing, derivatives, detrending and related filters change the spectral directions seen by partial least squares (PLS) or Ridge regression, but a full external search repeatedly refits nearly the same linear model. This paper studies the case where that search can be collapsed into one calibration step. For a strict linear preprocessing operator A acting on row spectra as XA^T, the transformed PLS cross-covariance satisfies (XA^T)^T Y = A X^T Y, and Ridge regression depends on the operator-induced kernel X A^T A X^T. These identities let a finite operator bank be screened inside the model while retaining original-wavelength coefficients, and the same identity extends to cheaply evaluated linear operator chains. Sample-adaptive or fitted corrections such as SNV, MSC, EMSC and ASLS are not strict linear; we prove the boundary and keep them as fold-local branches. The cohort has 61 regression and 17 classification rows, with a strict paired regression denominator of N=32 for the eight paper variants. There, AOM-PLS reaches median RMSEP ratios of 0.991/0.990 (simple) and 0.985/1.002 (best) against PLS-default/PLS-HPO, and AOM-Ridge reaches 0.974/0.984 (simple) and 0.918/0.966 (best) against Ridge-default/Ridge-HPO. The operator-adaptive classifier AOM-PLS-DA improves balanced accuracy by a median 0.159 on N=13 datasets (12/13 wins). The practical result is the runtime gap: PLS-HPO takes a median 710.81 s per run, whereas AOM-PLS takes 1.18-1.63 s -- 436 to 602 times less PLS fitting time. Linear operator-adaptive calibration thus gives prediction quality comparable to exhaustive preprocessing screening, with orders-of-magnitude less fitting time for PLS.

光谱分析模型加速线性建模预处理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。