提出双向分子-文本对齐框架,提升药物设计中的生成准确性。
RTMol: Rethinking Molecule-text Alignment in a Round-trip View
- 通过自监督往返学习统一分子描述与文本生成任务。
- 在多个大模型上实现双向对齐性能最高提升47%。
- 无需成对数据即可训练,适合化学文本生成场景。
将分子序列表示(如SMILES)与文本描述对齐,在药物发现、材料设计和自动化化学文献分析中至关重要。现有方法通常将分子描述生成(分子→文本)与基于文本的分子设计(文本→分子)视为独立任务,依赖监督微调或对比学习。这些方法存在三大局限:(i) 传统指标如BLEU侧重语言流畅性而非化学准确性;(ii) 训练数据常含化学不明确的描述和不完整规范;(iii) 双向生成独立优化导致不一致性。为此,我们提出RTMol,一种通过自监督往返学习统一分子描述与文本生成的双向对齐框架。该框架引入新型往返评估指标,支持无需成对分子-文本语料库的无监督分子描述训练。实验表明,RTMol在多种大模型上将双向对齐性能提升最高达47%,建立了一种有效的分子-文本联合理解与生成范式。
原文摘要 · Abstract (English)
Aligning molecular sequence representations (e.g., SMILES notations) with textual descriptions is critical for applications spanning drug discovery, materials design, and automated chemical literature analysis. Existing methodologies typically treat molecular captioning (molecule-to-text) and text-based molecular design (text-to-molecule) as separate tasks, relying on supervised fine-tuning or contrastive learning pipelines. These approaches face three key limitations: (i) conventional metrics like BLEU prioritize linguistic fluency over chemical accuracy, (ii) training datasets frequently contain chemically ambiguous narratives with incomplete specifications, and (iii) independent optimization of generation directions leads to bidirectional inconsistency. To address these issues, we propose RTMol, a bidirectional alignment framework that unifies molecular captioning and text-to-SMILES generation through self-supervised round-trip learning. The framework introduces novel round-trip evaluation metrics and enables unsupervised training for molecular captioning without requiring paired molecule-text corpora. Experiments demonstrate that RTMol enhances bidirectional alignment performance by up to 47% across various LLMs, establishing an effective paradigm for joint molecule-text understanding and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。