arXiv:2603.25062cs.LG2026-03

让分子生成模型对等价序列做出一致预测,提升生成质量与稳定性。

SIGMA: Semantic Identifier Grouping for Molecular Autoregression

  • 基于化学等价后缀三元组设计新目标函数,统一相同路径的中间状态
  • 在8个数据集组合中,6个实现分子质量评估指标显著降低(置信区间全负)
  • 无需改动模型结构,适用于生成与属性预测任务,适合分子建模研究者

自回归分子模型在不同分子序列化形式下仍需分配概率,但化学身份对序列顺序不变。等价序列可能对应相同分子却导致不一致的下一步预测。随机化序列虽能扩大覆盖范围,却无法揭示哪些中间决策应保持一致。本文提出SIGMA,一种基于化学认证等后缀三元组的密集后缀位置目标:两个等价历史、一个非等价历史与共享后缀。SIGMA在延续路径上对齐对应预令牌隐藏状态,将负样本分离至有限相对边界,且不改变语言模型目标、解码器及推理流程。在四种数据集上对比标准训练、随机序列训练和最后令牌对齐,在SMILES与SELFIES表示下,八个表示-数据组合中,有六个实现测试参考弗雷谢特化学网络距离显著下降,配对95%置信区间均低于零。逐位置分析显示状态一致性、化学区分度与下一步预测一致性均提升,同时保持分子间信息。在两个完整ZINC语料库块上,计算量匹配的消融实验表明,化学正确的状态对齐是有效关键成分。除生成外,SIGMA在所有六项分子属性基准上提升平均预测性能,并降低对等价分子序列化的敏感性。

原文摘要 · Abstract (English)

Autoregressive molecular models assign probability to molecular serializations even though chemical identity is invariant to serialization. Equivalent serializations can therefore represent a common molecular identity yet induce inconsistent next-token decisions. Randomized strings broaden exposure, but do not reveal which intermediate decisions should agree. We introduce SIGMA, a dense suffix-position objective built from chemically certified same-suffix triplets: two equivalent histories, one non-equivalent history, and a shared suffix. SIGMA aligns corresponding pre-token hidden states along the continuation and separates the negative to a finite relative margin, leaving the language-model objective, decoder, and inference procedure unchanged. We compare SIGMA with canonical training, randomized-serialization training, and last-token alignment across four datasets under SMILES and SELFIES. Across the eight representation-dataset blocks, SIGMA yields clear test-reference Frechet ChemNet Distance reductions in six: all four SELFIES domains and QM9 and ZINC under SMILES, with paired 95% confidence intervals below zero against every control. Position-wise analyses show improved state correspondence, chemical discrimination, and next-token agreement while preserving between-molecule information. On two full-corpus ZINC blocks, compute-matched ablations identify chemically correct state correspondence as the effective ingredient. Beyond generation, SIGMA improves mean predictive performance on all six molecular property benchmarks and reduces sensitivity to equivalent molecular serializations on every task.

分子生成自回归模型化学智能序列对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。