arXiv:2511.07433physics.chem-phcs.AI2025-11

用新采样法让量子化学数据生成成本降15-50倍,还能保持精度。

Benchmarking Simulacra AI's Quantum Accurate Synthetic Data Generation for Chemical Sciences

  • 提出RELAX采样法,结合波函数模型与变分蒙特卡洛算法。
  • 生成效率提升15-50倍,氨基酸级计算比传统CCSD快2-3倍。
  • 适合制药、材料等需大规模高精度数据的AI研发团队。

本文将Simulacra的合成数据生成流程与微软的先进方案在小到大体系的数据集上进行基准测试。通过分析能量质量、自相关时间与有效样本量,结果表明:搭配先进变分蒙特卡洛(VMC)采样算法的Simulacra大波函数模型(LWM)流程,使数据生成成本降低15-50倍,能量精度与现有方法相当;在氨基酸尺度上,相比传统CCSD方法提速2-3倍。该成果基于一种新颖且专有的采样策略——回火式朗之万自适应探索(RELAX),实现了低成本、大规模从头算数据集的构建,推动医药等领域中人工智能驱动的优化与发现进程。

原文摘要 · Abstract (English)

In this work, we benchmark \simulacra's synthetic data generation pipeline against a state-of-the-art Microsoft pipeline on a dataset of small to large systems. By analyzing the energy quality, autocorrelation times, and effective sample size, our findings show that Simulacra's Large Wavefunction Models (LWM) pipeline, paired with state-of-the-art Variational Monte Carlo (VMC) sampling algorithms, reduces data generation costs by 15-50x, while maintaining parity in energy accuracy, and 2-3x compared to traditional CCSD methods on the scale of amino acids. This enables the creation of affordable, large-scale \textit{ab-initio} datasets, accelerating AI-driven optimization and discovery in the pharmaceutical industry and beyond. Our improvements are based on a novel and proprietary sampling scheme called Replica Exchange with Langevin Adaptive eXploration (RELAX).

量子化学合成数据采样优化生物医药

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。