用自适应不确定性评估提升大模型在化学推理中的准确率
ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation
- 根据推理步骤位置动态调整不确定性值,精准定位知识盲区
- 在三个化学数据集上使大模型推理准确率显著提升
- 适合需要高可靠性的化学研究与药物发现领域
大语言模型(LLMs)在数学和编程任务中表现出色,但在化学问题上表现大幅下降。化学推理涉及长链条、复杂术语及专用符号系统,导致通用大模型易产生幻觉。现有方法难以有效融合化学专业知识,且不确定性估计无法精确定位错误环节。本文提出ChemAU框架,引入自适应不确定性估计机制,依据推理步骤在链中的位置分配不同不确定性值。该方法能精准识别化学知识缺口,并通过领域专用模型补充专业知识,修正原有错误推理链。在三个主流化学数据集上,对三种大模型的实验表明,ChemAU显著提升了推理准确率与不确定性估计精度。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are widely used across various scenarios due to their exceptional reasoning capabilities and natural language understanding. While LLMs demonstrate strong performance in tasks involving mathematics and coding, their effectiveness diminishes significantly when applied to chemistry-related problems. Chemistry problems typically involve long and complex reasoning steps, which contain specific terminology, including specialized symbol systems and complex nomenclature conventions. These characteristics often cause general LLMs to experience hallucinations during the reasoning process due to their lack of specific knowledge. However, existing methods are struggling to effectively leverage chemical expertise and formulas. Moreover, current uncertainty estimation methods, designed to mitigate potential reasoning errors, are unable to precisely identify specific steps or key knowledge. In this work, we propose a novel framework called ChemAU, which incorporates our adaptive uncertainty estimation method that applies different uncertainty values based on the position of reasoning steps within the whole reasoning chain. Leveraging this method, ChemAU identifies gaps in chemistry knowledge and precisely supplements chemical expertise with the specialized domain model, thereby correcting and updating the previously flawed reasoning chain. Our experiments with three popular LLMs across three chemistry datasets demonstrate that ChemAU significantly enhances both reasoning accuracy and uncertainty estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。