arXiv:2508.08401cs.CL2025-08被引 14

提升分子发现中大模型的推理能力与可解释性

Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery

  • 用先验引导的上下文蒸馏构建高质量推理数据集
  • 通过分子迭代适配策略显著提升生成效果
  • 适合药物设计与化学智能研发人员参考

大型语言模型(LLM)尤其是显式长链思维(Long-CoT)推理模型如DeepSeek-R1和QWQ,在常识推理和数学推断中表现优异。然而,这类模型在知识密集型领域如分子发现中的表现受限,效率低下。该任务需要对分子结构和化学原理有精确理解,而分子数据固有的复杂性及高质量专家标注的稀缺性增加了难度。为此,我们提出Mol-R1框架,旨在提升R1类显式长链思维模型在文本驱动分子生成中的可解释性与推理性能。首先,通过先验引导的上下文蒸馏(PRID)策略构建高质量推理数据集,生成受先验约束的推理轨迹。在此基础上,提出分子迭代适配(MoIA)训练策略,迭代结合监督微调(SFT)与强化策略优化(RPO),专门增强模型在分子发现任务中的推理能力。实验表明,Mol-R1在文本驱动分子推理生成任务中优于现有基线。

原文摘要 · Abstract (English)

Large language models (LLMs), especially Explicit Long Chain-of-Thought (CoT) reasoning models like DeepSeek-R1 and QWQ, have demonstrated powerful reasoning capabilities, achieving impressive performance in commonsense reasoning and mathematical inference. Despite their effectiveness, Long-CoT reasoning models are often criticized for their limited ability and low efficiency in knowledge-intensive domains such as molecule discovery. Success in this field requires a precise understanding of domain knowledge, including molecular structures and chemical principles, which is challenging due to the inherent complexity of molecular data and the scarcity of high-quality expert annotations. To bridge this gap, we introduce Mol-R1, a novel framework designed to improve explainability and reasoning performance of R1-like Explicit Long-CoT reasoning LLMs in text-based molecule generation. Our approach begins with a high-quality reasoning dataset curated through Prior Regulation via In-context Distillation (PRID), a dedicated distillation strategy to effectively generate paired reasoning traces guided by prior regulations. Building upon this, we introduce MoIA, Molecular Iterative Adaptation, a sophisticated training strategy that iteratively combines Supervised Fine-tuning (SFT) with Reinforced Policy Optimization (RPO), tailored to boost the reasoning performance of R1-like reasoning models for molecule discovery. Finally, we examine the performance of Mol-R1 in the text-based molecule reasoning generation task, showing superior performance against existing baselines.

分子生成长链推理模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。