arXiv:2605.29698cs.LGphysics.chem-ph2026-05

提出新评估框架,精准分离分子混合物的非理想行为预测误差。

A Systematic Evaluation of Molecular Mixture Behavior Prediction

论文配图:A Systematic Evaluation of Molecular Mixture Behavior Prediction
图 1 · 摘自论文原文
  • 分解混合物误差为纯组分与非理想相互作用两部分
  • 发现强绝对精度仍可能遗漏非理想行为,分子未见时性能骤降
  • 适合关注混合物建模泛化能力的研究者

分子性质预测的机器学习研究长期集中于纯物质,而实际应用多涉及存在分子间相互作用的混合物。近期工作扩充了混合物数据集,但评估仍主要依赖绝对准确率。然而,混合物中的绝对误差混杂了纯组分贡献与偏离理想混合的部分。本文提出一种评估框架,将混合物性质误差分解为纯组分与相互作用(非理想)成分。该框架结合泄漏感知的数据划分、理想混合基线和过量性质指标。为支持可复现基准测试,我们整理了七个匹配的纯物质与混合物物理化学性质数据集。在多个混合物性质任务和模型类型中,发现强绝对准确率可能掩盖对非理想混合行为的不良恢复,且在严格分子划分下性能显著下降。结果表明,向未见分子迁移是分子混合物机器学习的核心挑战,推动评估应超越单一绝对准确率。

原文摘要 · Abstract (English)

Machine learning for molecular property prediction has focused largely on pure compounds, even though many practical applications depend on mixtures with intermolecular interactions. Recent work has expanded the availability of mixture datasets, but evaluation still focuses mainly on absolute accuracy. However, absolute errors in mixtures conflate pure-component contributions with deviations from ideal mixing. We propose an evaluation framework that decomposes mixture-property error into pure-compound and interaction (non-ideal) components. The framework combines leakage-aware split protocols, ideal-mixture baselines, and excess-property metrics. To support reproducible benchmarking, we curate seven matched pure and mixture physicochemical property datasets. Across multiple mixture-property tasks and model families, we find that strong absolute accuracy can mask poor recovery of non-ideal mixture behavior, and that performance drops substantially under strict molecule splits. These results identify transfer to unseen molecules as a central challenge in molecular mixture machine learning and motivate evaluation beyond absolute accuracy alone.

分子混合物误差分解非理想行为评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。