arXiv:2607.23469physics.opticscs.AI2026-07

用强化学习加速光子器件设计,杜林DQN效果最佳。

When Every Simulation Counts: Value-Based Reinforcement Learning for Accelerated Photonics Inverse Design

论文配图:When Every Simulation Counts: Value-Based Reinforcement Learning for Accelerated Photonics Inverse Design
图 1 · 摘自论文原文
  • 对比七种价值学习方法,杜林DQN在有限仿真次数下表现最稳。
  • 相比初始设计,最优结构提升品质因子、减少波长误差64%。
  • 结果可复现,适合需高效仿真资源的科学优化任务。

光子晶体表面发射激光器(PCSEL)兼具高功率与窄发散角,但参数优化依赖昂贵的全波仿真。深度Q网络(DQN)可通过重用仿真数据指导设计迭代,但在严格仿真预算下,哪种价值学习机制仍可靠尚不明确。本文在统一目标、模拟器、83次调用预算及四组匹配初始化条件下,比较基线DQN与六种变体在七变量PCSEL设计中的表现。不仅分析终点性能,还考察样本效率、策略行为与物理响应,以区分学习收益与初始优势或探索跳跃。结果显示,唯有杜林DQN在全部四个种子上均实现改进;其选中结构使平均品质因子从12.5提升至20.7,波长误差降低64%,向上功率增加47%;相较基线DQN,在相同预算下达到更高平均性能。其他变体无一致提升:双DQN重复基线轨迹,Rainbow-lite虽有高潜力但严重依赖种子。研究确认杜林DQN为当前最可靠的配置,并提供可复现的算法增益归因框架。源码已公开于https://github.com/Longying-Wen/PCSEL-RL。

原文摘要 · Abstract (English)

Photonic-crystal surface-emitting lasers (PCSELs) can combine high-power operation with narrow-divergence surface emission, but optimizing coupled parameters requires costly full-wave simulations. Deep Q-network (DQN) optimization can reuse simulated transitions to guide edits, yet which value-learning mechanisms remain reliable under tight simulation budgets is unknown. We address this gap by comparing baseline DQN and six value-based variants for a seven-variable PCSEL design under a shared objective, simulator, 83-call budget, and four matched initializations. Beyond endpoints, we analyze sample efficiency, policy behavior, and physical response to separate learning gains from favorable starts or exploratory jumps. Dueling DQN is the only variant to improve all four seeds. Relative to the first evaluated designs, its selected structures increase the mean quality factor () from to (), reduce wavelength error by 64%, and increase upward power by 47%; compared with baseline DQN, they achieve a higher mean under the same budget. Other variants yield no consistent improvement; Double DQN reproduces baseline trajectories, while Rainbow-lite shows high upside but strong seed dependence. These results identify Dueling DQN as the most reliable configuration tested for simulation-budget-limited PCSEL inverse design and provide a reproducible framework for attributing algorithmic gains in scientific optimization. The source code is publicly available at https://github.com/Longying-Wen/PCSEL-RL.

强化学习光子设计仿真优化DQN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。