Logos模型通过逻辑推理与化学一致性结合,实现可解释的分子设计。
Logos: An evolvable reasoning engine for rational molecular design
- 分阶段训练:先学推理逻辑,再对齐分子结构,最后融入化学规则。
- 在多个数据集上达到高结构准确率和化学有效性,参数量仅为大模型的几分之一。
- 推理过程透明可查,适合需要可解释性的科研人员使用。
功能分子的发现与设计是化学、生物学和材料科学中的核心挑战。尽管机器学习在分子性质预测和候选生成方面取得进展,但现有模型往往在物理保真度与可解释推理之间难以兼顾,或在灵活性与化学合理性之间失衡,限制了其在真实科研流程中的可靠性。本文提出Logos,一个紧凑的分子推理模型,将多步逻辑推理与严格的化学一致性相结合。该模型采用分阶段训练策略:首先引入从分子描述到结构决策的显式推理示例,随后逐步对齐推理模式与分子表征;在最终阶段,直接将化学规则与守恒律纳入优化目标,引导模型生成化学有效的结果。在多个基准数据集上,Logos在结构准确性和化学有效性方面表现优异,性能媲美甚至超越更大规模的通用语言模型,而参数量仅为后者的极小部分。此外,在涉及多重潜在冲突约束的分子优化任务中,模型表现出稳定行为。通过显式展示中间推理步骤,Logos支持人类对生成结构的设计逻辑进行审查。这些结果表明,联合优化推理结构与物理一致性,为构建可靠且可解释的分子科学人工智能系统提供了可行路径,推动AI更深入地融入科学发现流程。
原文摘要 · Abstract (English)
The discovery and design of functional molecules remain central challenges across chemistry,biology, and materials science. While recent advances in machine learning have accelerated molecular property prediction and candidate generation, existing models tend to excel either in physical fidelity without transparent reasoning, or in flexible reasoning without guarantees of chemical validity. This imbalance limits the reliability of artificial intelligence systems in real scientific design workflows.Here we present Logos, a compact molecular reasoning model that integrates multi-step logical reasoning with strict chemical consistency. Logos is trained using a staged strategy that first exposes the model to explicit reasoning examples linking molecular descriptions to structural decisions, and then progressively aligns these reasoning patterns with molecular representations. In a final training phase, chemical rules and invariants are incorporated directly into the optimization objective, guiding the model toward chemically valid outputs. Across multiple benchmark datasets, Logos achieves strong performance in both structural accuracy and chemical validity, matching or surpassing substantially larger general-purpose language models while operating with a fraction of their parameters. Beyond benchmark evaluation, the model exhibits stable behaviour in molecular optimization tasks involving multiple, potentially conflicting constraints. By explicitly exposing intermediate reasoning steps, Logos enables human inspection and assessment of the design logic underlying each generated structure. These results indicate that jointly optimizing for reasoning structure and physical consistency offers a practical pathway toward reliable and interpretable AI systems for molecular science, supporting closer integration of artificial intelligence into scientific discovery processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。