arXiv:2607.29491cs.LGcs.AI2026-07

用世界模型减少量子电路搜索中的真实计算次数,提升效率。

DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search

  • 构建基于模型的强化学习框架,仅学习昂贵的后验反馈。
  • 在5个分子任务中平均降低1.6至10.6倍真实VQE调用次数。
  • 适合需要高效量子算法设计的研究者和工程师。

基于强化学习的量子架构搜索(RL-QAS)在扩展量子电路后反复优化变分量子本征值求解器(VQE),尽管电路构建与动作合法性是确定且已知的。我们提出DreamQAS,一种基于模型的强化学习框架,保留了精确的电路动态,仅学习昂贵的后验反馈。通过循环随机先验集成预测相对于经验能量前沿的无监督评分,并支持在合法电路上进行多步想象策略学习。基于排名的激活、不确定性感知的悲观性与截断,以及选择性真实VQE验证构成可靠性控制的学习闭环。在15,000次回合预算下,冻结评估条件下,DreamQAS在五个分子任务中的四个上实现了最低的平均冻结策略能量误差,一个任务中位居第二。在所有种子均达到精细误差目标时,其在四个任务中真实VQE调用次数减少1.6至2.0倍,在BeH2-8q任务中减少10.6倍。反事实动作排序效用在所有五项任务中提升,平均增加0.346,95%置信区间为[0.185, 0.507],而直接贪婪或束搜索无法恢复想象策略学习带来的收益。集成分歧也在所有三个测试任务中改善了风险覆盖能力。这些结果确立了一种以决策有用反馈为核心的量子架构搜索世界模型设计。

原文摘要 · Abstract (English)

Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extending a circuit, although circuit construction and action legality are deterministic and known. We introduce DreamQAS, a model-based RL framework that preserves these exact circuit dynamics and learns only the expensive post-VQE feedback. A recurrent randomized-prior ensemble predicts an oracle-free score relative to an empirical energy frontier and supports multi-step imagined policy learning over explicit legal circuits. Ranking-based activation, uncertainty-aware pessimism and truncation, and selective real-VQE verification form a reliability-controlled learning loop. Under a common 15,000-episode budget and frozen evaluation for the RL methods, DreamQAS has the lowest mean frozen-policy energy error on four of five molecular tasks and the second-lowest on one. At fine-error targets reached by all seeds of both methods, it uses 1.6x to 2.0x fewer real VQE calls on four tasks and 10.6x fewer on BeH2-8q. Counterfactual action-ranking utility increases across all five tasks, with a mean increase of 0.346 and a 95 percent confidence interval of [0.185, 0.507], while direct greedy and beam use of the same model does not recover the gains of imagined policy learning. Ensemble disagreement also improves risk-coverage over random rejection on all three probed tasks. These results establish a world-model design for QAS whose value lies in decision-useful feedback rather than exact energy prediction.

量子计算架构搜索强化学习量子算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。