提出新型自编码器,实现浑浊介质光谱解混无需校准
Bin Latent Transformer (BiLT): A shift-invariant autoencoder for calibration-free spectral unmixing of turbid media

- 用交叉注意力扫描替代全连接编码器,摆脱波长位置依赖
- 测试集上吸收/散射系数预测相关系数达0.979和0.975
- 对仪器漂移和硬件更换有强鲁棒性,适合实际应用
从积分球测量中准确恢复组分光学特性是制药分析、食品科学和生物医学诊断中的核心挑战。神经网络自编码器可无先验知识提取各组分的光谱吸收与散射系数,但其全连接编码器将特征绑定于绝对波长索引,导致光谱仪校准漂移或硬件更换时精度下降。本文提出Bin Latent Transformer(BiLT)-自编码器,用跨注意力扫描器替代密集编码器:16个可学习探针向卷积特征图查询,独立于绝对波长位置聚合形态学光谱信息。结合物理约束线性解码器(强制吸收/散射分离)与三阶段课程增强策略构建完整架构。在液体模拟物基准测试中(含脂肪乳及两种油墨吸收体;496样本),模型在保留测试谱上对μ_a(λ)和μ_s'(λ)的预测相关系数分别为0.979和0.975,在±10个光谱带的全漂移范围内仍保持μ_a的R² > 0.90,μ_s'的R² ≈ 0.99。该模型在未重新训练条件下即可泛化至具有更宽仪器线型(≈24nm FWHM)的模拟光谱仪,对两通道仍保持R² ≈ 0.96和0.974。注意力图分析揭示一种物理可解释的双组件探针策略:在吸收边波长处设置稀疏锚点探针,同时在高透射长波段区域采用扩散式、信噪比驱动的探针集合,并在噪声下动态补充探针以实现隐式光谱平均。
原文摘要 · Abstract (English)
The accurate recovery of constituent-level optical properties from integrating sphere measurements is a central analytical challenge in pharmaceutical analysis, food science, and biomedical diagnostics. Neural network autoencoders can extract spectrally resolved absorption and scattering coefficients for each constituent without prior knowledge, but their fully connected encoders bind learned features to absolute wavelength indices, causing accuracy loss under spectrometer calibration drift or hardware exchange. This work introduces the Bin Latent Transformer (BiLT)-Autoencoder, in which the dense encoder is replaced by a cross-attention scanner: 16 learnable probe vectors query a convolutional feature map, aggregating morphological spectral information independently of absolute wavelength position. A physics-constrained linear decoder with enforced absorption/scattering separation and a three-phase curriculum augmentation strategy complete the architecture. On a liquid phantom benchmark (intralipid and two ink absorbers; 496 samples), the model achieves $R^2 = 0.979$ and $0.975$ for $μ_a(λ)$ and $μ_s'(λ)$, respectively, on held-out test spectra, maintaining $R^2 > 0.90$ for $μ_a$ and $R^2 \approx 0.99$ for $μ_s'$ across the full tested shift range of $\pm 10$ spectral bands. The model generalises to a simulated spectrometer with a broader instrument line shape (${\approx}24$nm FWHM) without retraining, retaining $R^2 \approx 0.96$ and $0.974$ for the two channels. Attention map analysis reveals a physically interpretable two-component probe strategy: sparse anchor probes at absorption-edge wavelengths combined with a diffuse, SNR-driven ensemble at the high-transmittance long-wavelength region, which recruits additional probes dynamically under noise to provide implicit spectral averaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。