通过双向验证训练,让化学大模型更准确地还原文本描述的分子结构。
Round-trip Reinforcement Learning: Self-Consistent Training for Better Chemical LLMs
- 用双向转换成功作为奖励信号,驱动模型自我优化一致性。
- 在多种数据环境下显著提升模型性能与结构还原准确率。
- 适合追求高可靠性化学生成模型的研究者和开发者。
大语言模型(LLMs)正成为计算化学的通用基础模型,可处理反应预测与逆合成等双向任务。然而,这些模型常缺乏双向一致性:例如,一个先进的化学LLM能准确描述分子结构,却无法从自动生成的文本中还原原始结构。这种不一致表明模型仅学习了单向记忆而非灵活掌握。近期研究证实,模型的双向一致性与其主任务表现呈强相关性,这使一致性成为直接优化目标。本文提出往返强化学习(RTRL),通过往返转换的成功作为奖励信号,训练模型提升一致性。进一步提出迭代版本,正向与反向映射交替训练,形成自增强循环,具有高度数据效率,尤其适用于化学领域海量无标注数据。实验表明,RTRL在监督、自监督及合成数据设置下均显著优于强基线模型,显著提升性能与一致性。该工作证明,双向一致性不仅是理想属性,更是可训练目标,为构建更鲁棒、可靠的化学基础模型提供了新路径。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are emerging as versatile foundation models for computational chemistry, handling bidirectional tasks like reaction prediction and retrosynthesis. However, these models often lack round-trip consistency. For instance, a state-of-the-art chemical LLM may successfully caption a molecule, yet be unable to accurately reconstruct the original structure from its own generated text. This inconsistency suggests that models are learning unidirectional memorization rather than flexible mastery. Indeed, recent work has demonstrated a strong correlation between a model's round-trip consistency and its performance on the primary tasks. This strong correlation reframes consistency into a direct target for model improvement. We therefore introduce Round-Trip Reinforcement Learning (RTRL), a novel framework that trains a model to improve its consistency by using the success of a round-trip transformation as a reward signal. We further propose an iterative variant where forward and reverse mappings alternately train each other in a self-improvement loop, a process that is highly data-efficient and notably effective with the massive amount of unlabelled data common in chemistry. Experiments demonstrate that RTRL significantly \textbf{boosts performance and consistency} over strong baselines across supervised, self-supervised, and synthetic data regimes. This work shows that round-trip consistency is not just a desirable property but a trainable objective, offering a new path toward more robust and reliable foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。