arXiv:2605.29329q-bio.QMcs.LG2026-05

用混合向量模型实现可精确求解的共聚物逆向设计。

Mixing Vector Model for Copolymer Inference via Mixed Integer Linear Programming

论文配图:Mixing Vector Model for Copolymer Inference via Mixed Integer Linear Programming
图 1 · 摘自论文原文
  • 提出混合向量模型,用单体描述加权组合表示共聚物特征。
  • 在10个数据集上测试,9个达R²>0.7,6个达R²>0.9。
  • 支持多单体逆向设计,且优化问题仍可高效求解。

本文将一种新型两阶段分子反向设计框架mol-infer拓展至共聚物领域,引入简化的混合向量(MV)模型。该模型将共聚物特征表示为单体描述符的凸组合,权重由单体比例决定,无需序列信息,天然适配基于混合整数线性规划(MILP)的设计。利用神经网络、简化二次多元线性回归和随机森林构建多个共聚物性质预测模型。在多个物理化学性质数据集上,该表示方法表现优异:10个数据集中有9个测试R²超过0.7,6个超过0.9。同时,在给定混合比例下构建多单体逆向设计问题,证明其生成的MILP实例对三单体情形仍具可解性。最后通过外部一致性验证,重新计算候选结构属性并与原模型预测值对比,确认结果可靠性。该框架为两层模型下的共聚物模型级精确逆向设计提供了可行起点。

原文摘要 · Abstract (English)

A novel two-phase molecule inference framework, mol-infer, has recently been developed to infer chemical graphs with prescribed abstract structures and desired property values through mixed integer linear programming (MILP) under the two-layered model, with guaranteed optimality and exactness relative to the given learned prediction function and structural constraints. In this study, we extend this framework to copolymers by introducing a simple feature representation, called the mixing vector (MV) model. In the proposed model, a copolymer feature vector is represented as a convex combination of MILP-tractable monomer descriptors weighted by the mixing ratio of the constituent monomers. This representation does not require explicit sequence-class information and is therefore naturally compatible with MILP-based inverse design. Under this model, we construct prediction functions for several copolymer property datasets using artificial neural networks, reduced quadratic multiple linear regression, and random forests. The proposed representation achieves practically useful predictive performance across multiple physicochemical property datasets; in particular, the best test R^2 score exceeds 0.7 for nine of the ten datasets and exceeds 0.9 for six datasets. We also formulate a multi-monomer inverse-design problem under the MV representation with a prescribed mixing ratio and show that the resulting MILP instances remain tractable, even for three-monomer settings. Finally, we perform an external consistency check by re-evaluating the inferred candidates and comparing the re-computed property values with those predicted by the learned model. Overall, the proposed framework gives a tractable first step toward model-level exact inverse design of copolymers under the two-layered model.

共聚物逆向设计MILP机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。