arXiv:2607.20935cs.CEcs.CL2026-07被引 1

化学大模型的推理过程像会出错的草稿纸,常编造分子结构。

Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

  • 用分子草图(如SMILES片段)辅助推理,但易产生虚构结构
  • 正确答案常伴随虚假结构描述,且草图对生成有因果影响
  • 适合关注模型推理可信度的研究者和开发者

化学推理语言模型本应通过忠实的思维链(CoT)推导分子答案,但在四个模型家族与十二项化学任务中,幻觉现象普遍且与答案正确性解耦:正确答案常伴随真实分子中不存在的虚构结构陈述。尽管如此,推理轨迹并非无意义。归因分析表明,模型以特定形式实现共享的草稿功能:Chem-R 和 ether-0 依赖碎片化 SMILES 草图,ChemDFM-R 则强调骨架、位置与命名线索。值得注意的是,扰动 Chem-R 的 SMILES 草图会显著降低生成质量,表明结构草图即使在文字描述失真时仍具因果作用。结果表明,化学 CoT 既非忠实解释,也非单纯事后合理化,而是一种易出错的分子草稿。这一发现警示不应将 CoT 视为推理真实的直接证据,并呼吁超越仅评价答案的全过程监督。

原文摘要 · Abstract (English)

Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning model families and twelve chemistry tasks, hallucination is widespread and largely decoupled from answer correctness: correct answers often coexist with fabricated structural claims absent from the relevant molecules. Yet this does not make the reasoning trace computationally irrelevant. Attribution analyses suggest a shared scratchpad function expressed in model-specific forms: Chem-R and ether-0 rely on fragmented SMILES drafts, whereas ChemDFM-R emphasizes scaffold, positional, and naming cues. Notably, perturbing Chem-R's SMILES sketches degrades generation, showing that structural drafts can be causally load-bearing even when verbal structural claims are largely inert. Together, these results show that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization, but a hallucination-prone molecular scratchpad. This finding cautions against treating CoT as direct evidence of faithful reasoning and motivates process-level supervision beyond answer-only evaluation.

化学推理思维链幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。