构建离子液体碳捕集领域数据集,评测大模型科学推理能力。
Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research
- 构建5920条专家标注数据,覆盖不同难度的离子液体碳捕集问题。
- 小规模通用大模型仅具基础知识,缺乏高级科研推理能力。
- 为低碳研究与大模型协同进化提供可落地的评估工具。
大型语言模型(LLMs)在多个领域展现出卓越的通用知识与推理能力,但在化学与生物工程(CBE)等专业领域的表现仍待深入探索。为此,我们开展了一项针对离子液体(ILs)用于碳捕集这一应对全球变暖新兴技术的全面实证分析。我们构建并发布了一个包含5,920个示例的专家标注数据集,用于评估大模型在该领域的推理能力。数据集涵盖不同难度层级,平衡语言复杂度与专业领域知识。基于此,我们对三个参数少于100亿的开源大模型进行了评估,结果表明:尽管小型通用大模型具备基本的离子液体知识,但缺乏应用于高级科研任务所需的专门推理技能。在此基础上,我们探讨了提升大模型在碳捕集研究中实用性的策略。考虑到大模型自身的高碳足迹,将其发展与离子液体研究相结合,有望推动两领域协同发展,助力实现2050年碳中和目标。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated exceptional performance in general knowledge and reasoning tasks across various domains. However, their effectiveness in specialized scientific fields like Chemical and Biological Engineering (CBE) remains underexplored. Addressing this gap requires robust evaluation benchmarks that assess both knowledge and reasoning capabilities in these niche areas, which are currently lacking. To bridge this divide, we present a comprehensive empirical analysis of LLM reasoning capabilities in CBE, with a focus on Ionic Liquids (ILs) for carbon sequestration - an emerging solution for mitigating global warming. We develop and release an expert - curated dataset of 5,920 examples designed to benchmark LLMs' reasoning in this domain. The dataset incorporates varying levels of difficulty, balancing linguistic complexity and domain-specific knowledge. Using this dataset, we evaluate three open-source LLMs with fewer than 10 billion parameters. Our findings reveal that while smaller general-purpose LLMs exhibit basic knowledge of ILs, they lack the specialized reasoning skills necessary for advanced applications. Building on these results, we discuss strategies to enhance the utility of LLMs for carbon capture research, particularly using ILs. Given the significant carbon footprint of LLMs, aligning their development with IL research presents a unique opportunity to foster mutual progress in both fields and advance global efforts toward achieving carbon neutrality by 2050.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。