用大模型+文献知识,让红外光谱分析在小数据下也能准且快。
LUMIR: an LLM-Driven Unified Agent Framework for Multi-task Infrared Spectroscopy Reasoning
- 基于文献挖掘构建知识库,自动处理光谱数据并提取特征
- 在多个数据集上表现超越传统模型,尤其在低数据量时优势明显
- 适合科研与工业界做快速、自动化光谱分析的团队使用
红外光谱可实现化学与材料属性的快速无损分析,但高维信号与重叠峰阻碍了传统化学生物计量方法的应用。大型语言模型(LLMs)具备强泛化与推理能力,为自动化光谱解读带来新机遇,但在该领域应用仍不充分。本文提出LUMIR(LLM驱动的统一智能体框架,用于多任务红外光谱推理),一个面向低数据条件下精准红外光谱分析的代理框架。LUMIR整合结构化文献知识库、自动化预处理、特征提取与预测建模于一体,通过挖掘同行评审光谱研究,识别已验证的预处理与特征提取策略,将光谱转换为低维表示,并采用少样本提示完成分类、回归与异常检测任务。框架在多个数据集上验证,包括公开的Milk近红外数据集、中药材、陈皮(CRP)不同存储时长样本、工业废水COD数据集,以及Tecator和Corn两个公开基准。在各项任务中,LUMIR性能达到或优于现有机器学习与深度学习模型,尤其在资源受限场景下表现突出。本工作表明,结合结构化文献指导与少样本学习,可实现鲁棒、可扩展、自动化的光谱解读。LUMIR建立了一种将LLMs应用于红外光谱的新范式,以极少标注数据实现高精度,广泛适用于科学与工业领域。
原文摘要 · Abstract (English)
Infrared spectroscopy enables rapid, non destructive analysis of chemical and material properties, yet high dimensional signals and overlapping bands hinder conventional chemometric methods. Large language models (LLMs), with strong generalization and reasoning capabilities, offer new opportunities for automated spectral interpretation, but their potential in this domain remains largely untapped. This study introduces LUMIR (LLM-driven Unified agent framework for Multi-task Infrared spectroscopy Reasoning), an agent based framework designed to achieve accurate infrared spectral analysis under low data conditions. LUMIR integrates a structured literature knowledge base, automated preprocessing, feature extraction, and predictive modeling into a unified pipeline. By mining peer reviewed spectroscopy studies, it identifies validated preprocessing and feature derivation strategies, transforms spectra into low dimensional representations, and applies few-shot prompts for classification, regression, and anomaly detection. The framework was validated on diverse datasets, including the publicly available Milk near-infrared dataset, Chinese medicinal herbs, Citri Reticulatae Pericarpium(CRP) with different storage durations, an industrial wastewater COD dataset, and two additional public benchmarks, Tecator and Corn. Across these tasks, LUMIR achieved performance comparable to or surpassing established machine learning and deep learning models, particularly in resource limited settings. This work demonstrates that combining structured literature guidance with few-shot learning enables robust, scalable, and automated spectral interpretation. LUMIR establishes a new paradigm for applying LLMs to infrared spectroscopy, offering high accuracy with minimal labeled data and broad applicability across scientific and industrial domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。