arXiv:2604.04717cs.LGcond-mat.mtrl-sci2026-04被引 1

高维光谱数据让模型误判化学差异,即使无真实区别也能完美分类。

The Infinite-Dimensional Nature of Spectroscopy and Why Models Succeed, Fail, and Mislead

  • 利用高维空间理论解释模型为何能分清微小噪声差异
  • 实验显示无化学区别的光谱也能达99%以上准确率
  • 提醒研究者警惕特征重要性图的误导性

机器学习模型在光谱分类任务中表现惊人,却常无法证明其使用了化学有意义的特征。现有研究多关注预处理、噪声敏感性和模型复杂度,但缺乏统一解释。本文基于Feldman-Hajek定理与测度集中理论,揭示这些现象源于光谱数据固有的高维特性。即使由噪声、归一化或仪器误差引发的微小分布差异,在高维空间中也可能完全可分。通过合成与真实荧光光谱的系列实验,我们展示了模型在无化学差异时仍可达近完美准确率,并说明特征重要性图可能突出光谱无关区域。本文提供严谨理论框架,实验验证该效应,并提出光谱建模与解释的实践建议。

原文摘要 · Abstract (English)

Machine learning (ML) models have achieved strikingly high accuracies in spectroscopic classification tasks, often without a clear proof that those models used chemically meaningful features. Existing studies have linked these results to data preprocessing choices, noise sensitivity, and model complexity, but no unifying explanation is available so far. In this work, we show that these phenomena arise naturally from the intrinsic high dimensionality of spectral data. Using a theoretical analysis grounded in the Feldman-Hajek theorem and the concentration of measure, we show that even infinitesimal distributional differences, caused by noise, normalisation, or instrumental artefacts, may become perfectly separable in high-dimensional spaces. Through a series of specific experiments on synthetic and real fluorescence spectra, we illustrate how models can achieve near-perfect accuracy even when chemical distinctions are absent, and why feature-importance maps may highlight spectrally irrelevant regions. We provide a rigorous theoretical framework, confirm the effect experimentally, and conclude with practical recommendations for building and interpreting ML models in spectroscopy.

光谱分析高维数据机器学习模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。