arXiv:2606.07656physics.chem-phcs.CE2026-06

构建新溶剂溶解度基准SC3,解决旧数据集偏差问题。

SC3: The Multi-Solvent Solubility Challenge and Benchmark

论文配图:SC3: The Multi-Solvent Solubility Challenge and Benchmark
图 1 · 摘自论文原文
  • 基于BigSolDB v2.1建立可复现的清洗流程,获10万+测量数据
  • 重新校准随机误差下界为0.106 log S,比旧值收紧6倍
  • 引入分级共识与多溶剂评估指标,适合模型可靠性诊断

溶解度预测是计算化学的标准基准,但现有多溶剂模型虽声称逼近实验噪声上限(即随机误差极限),仍不可靠。我们指出该差距部分源于人为偏差:已有基准在数据清洗、评估指标(计数加权RMSE)及对0.6-0.8 log S间实验室差异的误判上存在缺陷。为此,我们提出SC3——基于BigSolDB v2.1的多溶剂溶解度基准,包含三项贡献:(i) 可复现的数据清洗流程,生成101,535条测量值,覆盖1,327种溶质与206种溶剂,重新校准随机误差下界为0.106 log S(约比传统值紧6倍);(ii) 分级黄金/白银/青铜共识层级,每点带标准差,三组漏检检查划分,以及多溶剂评估套件(PS-RMSE、Z-RMSE);(iii) 对六个家族共31个模型的基准测试,最佳青铜级PS-RMSE为随机误差极限的5倍,且所有深度模型均未填补此差距。后续分析包括数据缩放、量子化学溶剂化能迁移、特征级归因,证明校准后的点级不确定性是超越点预测诊断的可复用基础设施。

原文摘要 · Abstract (English)

Solubility prediction is a standard benchmark in computational chemistry, yet multi-solvent models which reportedly approach the experimental-noise ceiling (i.e. the aleatoric limit) are not yet reliable enough to be deployed. We argue that this gap is partly artefactual: published benchmarks differ in curation policies, evaluate on count-weighted RMSE that hides failure on tail-heavy solvent distributions, and treat the widely cited 0.6-0.8 log S inter-laboratory figure as the aleatoric ceiling even though it reflects worst-case, not expected, disagreement. We introduce SC3, a multi-solvent solubility benchmark built on BigSolDB v2.1 with three contributions: (i) a reproducible curation pipeline yielding 101,535 measurements over 1,327 solutes and 206 solvents, with a recalibrated aleatoric floor of 0.106 log S-roughly 6 times tighter than the conventional figure; (ii) nested Gold/Silver/Bronze consensus tiers with per-point standard deviation, three leakage-checked splits, and a multi-solvent metric suite (PS-RMSE, Z-RMSE); and (iii) a 31-model benchmark across six families, whose best Bronze PS-RMSE sits at 5 times the aleatoric limit, and we observe this is a gap unclosed by any deep alternative tested. We perform three follow-on analyses: data scaling, transfer from quantum-chemistry solvation energies, and feature-level attribution, which demonstrates that calibrated per-point uncertainty is a reusable infrastructure for diagnosis beyond point prediction.

溶剂溶解度基准测试不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。