arXiv:2606.17077physics.chem-phcs.AI2026-06

用量子辅助生成稀有pKa分子,解决数据稀缺难题。

Comprehensive pKa Data Augmentation from Limited Real Data through an Engineered Models-Quantum Framework

论文配图:Comprehensive pKa Data Augmentation from Limited Real Data through an Engineered Models-Quantum Framework
图 1 · 摘自论文原文
  • 构建机器学习模型预测大量未标注分子的pKa值。
  • 发现未标注数据中极端pKa值样本极度稀少。
  • 首次在物理量子退火机上实现稀有pKa分子高效生成。

质子解离常数(pKa)对功能分子发现和分子建模至关重要。基于iBonD——目前最大的实验pKa数据库,我们与合作者开发了基于机器学习的经验预测方法和高精度能量计算方法。尽管已有基础,高质量pKa数据的快速扩充仍受制于根本性瓶颈。本研究利用一系列高度优化的机器学习模型,在大量未标注分子数据集上进行大规模回归预测。结果显示,未标注分子数据的特征分布使pKa分布近似正态,极端区域样本极其稀少。虽该扩充对提升数据可用性和建模性能极具价值,但对发现具有广谱pKa特性的分子仍显不足。为此,我们探索从庞大化学空间中靶向生成稀有pKa分子。鉴于传统连续潜在空间VAE-RNN方法稳定性差且难以有效补充稀疏数据,我们设计并实现了量子辅助的稀有pKa分子生成。可行性已在模拟量子退火器上验证,并在物理相干伊辛机(CIMs)上实现更优的极端值采样。

原文摘要 · Abstract (English)

Proton dissociation constants (pKa) are critical for functional molecule discovery and molecular modeling. Building on iBonD, the largest experimental pKa database established, we and other researchers have developed several methods including machine-learning-based empirical prediction and high-accuracy energy calculations. Despite this foundation, the rapid augmentation of high-quality pKa data remains fundamentally constrained. As part of this work, we performed large-scale regression-based pKa prediction on unlabeled molecular datasets using a collection of extensively optimized machine-learning models. The results indicate that, since the feature distributions of unlabeled molecular datasets, the pKa data distribution approximates normality, with extreme scarcity of tail-region samples. Although such augmentation is highly valuable for improving overall data availability and predictive modeling, it remains insufficient for efficiently discovering molecules with broad-spectrum pKa properties. To address this, we explore the targeted generation of molecules with sparse pKa properties from the vast chemical space. Given that traditional continuous latent space VAE-RNN methods for molecular generation suffer from insufficient stability and fail to demonstrate clear advantages in complementing sparse data, we design and implement a quantum-assisted sparse-pKa molecular generation. Feasibility is validated on a simulated quantum annealer, and superior extreme-value sampling is further achieved on physical coherent Ising machines (CIMs). (to be continued)

pKa预测量子生成分子设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。