arXiv:2504.12314cs.CLcs.AI2025-04被引 8

提出新评估方法,精准检测化学大模型的虚构错误

How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension

  • 设计可计算的Mol-Hallu指标,衡量生成文本与真实分子属性的科学蕴含关系
  • 发现PubChem数据集存在知识捷径问题,导致模型频繁虚构分子性质
  • 提出后处理修正方案HRPP,有效降低各类分子大模型的幻觉率

大语言模型在分子理解任务中广泛应用,但存在幻觉问题,影响药物设计准确性。本文首次分析了分子理解任务中幻觉的成因,发现源于PubChem数据集的知识捷径现象。为高效评估幻觉,提出全新自由格式度量标准Mol-Hallu,基于生成文本与真实分子属性间的科学蕴含关系量化幻觉程度。利用该指标,重新评估多种分子大模型的幻觉水平。进一步提出幻觉抑制后处理阶段(HRPP),在decoder-only和encoder-decoder型分子大模型上均验证其有效性。研究结果为提升科学领域大模型可靠性提供了关键洞见。

原文摘要 · Abstract (English)

Large language models are increasingly used in scientific domains, especially for molecular understanding and analysis. However, existing models are affected by hallucination issues, resulting in errors in drug design and utilization. In this paper, we first analyze the sources of hallucination in LLMs for molecular comprehension tasks, specifically the knowledge shortcut phenomenon observed in the PubChem dataset. To evaluate hallucination in molecular comprehension tasks with computational efficiency, we introduce \textbf{Mol-Hallu}, a novel free-form evaluation metric that quantifies the degree of hallucination based on the scientific entailment relationship between generated text and actual molecular properties. Utilizing the Mol-Hallu metric, we reassess and analyze the extent of hallucination in various LLMs performing molecular comprehension tasks. Furthermore, the Hallucination Reduction Post-processing stage~(HRPP) is proposed to alleviate molecular hallucinations, Experiments show the effectiveness of HRPP on decoder-only and encoder-decoder molecular LLMs. Our findings provide critical insights into mitigating hallucination and improving the reliability of LLMs in scientific applications.

分子生成幻觉检测大模型评估科学计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。