用可验证的约束引导强化学习,加速中微子模型发现
CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery

- 通过符号推理生成失败原因证书,转化为可复用的约束条件
- 在超10^26种模型空间中,有效率提升至6.33倍,候选数减少4倍
- 提取可解释规则,再利用可进一步提升发现效率
许多科学发现问题需在复杂约束下搜索组合假设空间。强化学习虽具潜力,但现有方法依赖单一奖励信号,难以提供失败原因信息,导致智能体反复探索无效区域。本文提出认证驱动强化学习(CDRL),利用符号推理工具生成结构化反馈:当候选解违反领域约束时,工具会输出识别失败动作的证书。CDRL将这些证书转化为可复用的约束,排除无效解类别,引导探索向合法区域聚焦。我们在理论粒子物理中的中微子味模型发现任务上评估了该方法,假设空间超过10^26种可能模型,并与此前最优的强化学习方法对比。在三个理论空间中,CDRL实现最高1.95倍的有效模型率、最高6.33倍的中微子模型率,且评估候选数最多减少4倍。进一步通过后处理决策树框架从搜索轨迹中提取40条可解释规则,将其作为软约束重用后,在所有三个空间中有效模型率提升达2倍,中微子模型发现效率提升达3倍。结果表明,CDRL能发现组合搜索空间中的可复用结构,为科学模型发现提供通用框架。
原文摘要 · Abstract (English)
Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain constraints. Reinforcement learning (RL) offers a promising approach, but existing methods rely on scalar rewards that provide limited information about why candidate solutions fail, leading agents to repeatedly explore invalid regions. We introduce Certification-Driven Reinforcement Learning (CDRL), a framework that leverages structured feedback from symbolic reasoning tools. When a candidate violates domain constraints, these tools produce certificates identifying the actions responsible for failure. CDRL converts these certificates into reusable constraints that eliminate classes of invalid solutions and guide exploration toward valid regions. We evaluate CDRL on neutrino flavor model discovery in theoretical particle physics, where the hypothesis space exceeds $10^{26}$ possible models, and compare it with the state-of-the-art RL approach previously used for this task. Across three theory spaces, CDRL achieves up to 1.95$\times$ higher valid model rates and up to 6.33$\times$ higher neutrino model rates while evaluating up to 4$\times$ fewer candidates. We further extract 40 interpretable rules from search trajectories using a post-hoc decision-tree framework and show that reusing them as soft constraints yields gains of up to 2$\times$ in valid model rates and 3$\times$ in neutrino model discovery across all three theory spaces. These results suggest that CDRL uncovers reusable structure in combinatorial search spaces and provides a general framework for scientific model discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。