用张量分解+主动学习,精准预测混合物扩散系数随温度变化。
Prediction of Diffusion Coefficients in Mixtures with Tensor Completion
- 基于张量分解的混合模型,融合多温度实验数据与先验模型。
- 在268K至378K范围内预测误差显著低于传统模型。
- 结合核磁共振主动采样,有效扩充实验数据提升精度。
预测混合物中无限稀释状态下的扩散系数对众多应用至关重要,但实验数据稀缺,机器学习提供了替代传统半经验模型的潜力。矩阵补全方法在二元混合物的热物理性质预测中表现良好,但仅限单温度预测,且精度依赖高质量实验数据。本文提出一种混合张量补全方法(TCM),采用塔克分解,在298K、313K和333K的实验数据基础上联合训练,并引入半经验SEGWE模型作为贝叶斯框架中的先验知识。该方法可线性外推至268K至378K区间,预测精度显著优于现有模型。为进一步提升性能,通过主动学习策略,利用脉冲场梯度核磁共振(PFG NMR)测量了19个溶质-溶剂体系在298K、313K和333K下的扩散系数,扩充实验数据库后,显著提升了TCM的预测能力。结果表明,数据高效机器学习与自适应实验相结合,有望推动传输性质建模的发展。
原文摘要 · Abstract (English)
Predicting diffusion coefficients in mixtures is crucial for many applications, as experimental data remain scarce, and machine learning (ML) offers promising alternatives to established semi-empirical models. Among ML models, matrix completion methods (MCMs) have proven effective in predicting thermophysical properties, including diffusion coefficients in binary mixtures. However, MCMs are restricted to single-temperature predictions, and their accuracy depends strongly on the availability of high-quality experimental data for each temperature of interest. In this work, we address this challenge by presenting a hybrid tensor completion method (TCM) for predicting temperature-dependent diffusion coefficients at infinite dilution in binary mixtures. The TCM employs a Tucker decomposition and is jointly trained on experimental data for diffusion coefficients at infinite dilution in binary systems at 298 K, 313 K, and 333 K. Predictions from the semi-empirical SEGWE model serve as prior knowledge within a Bayesian training framework. The TCM then extrapolates linearly to any temperature between 268 K and 378 K, achieving markedly improved prediction accuracy compared to established models across all studied temperatures. To further enhance predictive performance, the experimental database was expanded using active learning (AL) strategies for targeted acquisition of new diffusion data by pulsed-field gradient (PFG) NMR measurements. Diffusion coefficients at infinite dilution in 19 solute + solvent systems were measured at 298 K, 313 K, and 333 K. Incorporating these results yields a substantial improvement in the TCM's predictive accuracy. These findings highlight the potential of combining data-efficient ML methods with adaptive experimentation to advance predictive modeling of transport properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。