用强化学习直接训练光学处理器,省去建模步骤,更快更稳地实现复杂光路调控。
Model-free Optical Processors using In Situ Reinforcement Learning with Proximal Policy Optimization
- 采用PPO算法在真实系统上直接优化光学结构,无需预先建模
- 在聚焦、全息成像等任务中收敛速度提升,性能优于传统方法
- 适合需要快速适应硬件缺陷的光学系统开发人员
光学计算有望实现高速、低功耗的信息处理,其中衍射光学网络作为灵活平台可用于实现特定任务的变换。然而,衍射层的有效优化与对齐面临挑战,因难以准确建模包含硬件缺陷、噪声和错位的物理系统。现有原位优化方法虽可在真实系统上直接训练而无需显式建模,但常因测量数据利用效率低导致收敛慢、性能不稳定。本文提出一种基于近端策略优化(PPO)的无模型强化学习方法,用于衍射光学处理器的原位训练。PPO能高效重用原位测量数据,并通过限制策略更新保证更稳定、更快的收敛。我们在多个原位学习任务中实验验证了该方法,包括通过随机散射体实现靶向能量聚焦、全息图像生成、像差校正和光学图像分类,各项任务均表现出更优的收敛性和性能。该策略直接作用于物理系统,自然涵盖未知的真实世界缺陷,无需先验系统知识或建模。通过在实际实验约束下实现更快更准的训练,这一原位强化学习方法可为受复杂反馈动态支配的各类光学与物理系统提供可扩展的训练框架。
原文摘要 · Abstract (English)
Optical computing holds promise for high-speed, energy-efficient information processing, with diffractive optical networks emerging as a flexible platform for implementing task-specific transformations. A challenge, however, is the effective optimization and alignment of the diffractive layers, which is hindered by the difficulty of accurately modeling physical systems with their inherent hardware imperfections, noise, and misalignments. While existing in situ optimization methods offer the advantage of direct training on the physical system without explicit system modeling, they are often limited by slow convergence and unstable performance due to inefficient use of limited measurement data. Here, we introduce a model-free reinforcement learning approach utilizing Proximal Policy Optimization (PPO) for the in situ training of diffractive optical processors. PPO efficiently reuses in situ measurement data and constrains policy updates to ensure more stable and faster convergence. We experimentally validated our method across a range of in situ learning tasks, including targeted energy focusing through a random diffuser, holographic image generation, aberration correction, and optical image classification, demonstrating in each task better convergence and performance. Our strategy operates directly on the physical system and naturally accounts for unknown real-world imperfections, eliminating the need for prior system knowledge or modeling. By enabling faster and more accurate training under realistic experimental constraints, this in situ reinforcement learning approach could offer a scalable framework for various optical and physical systems governed by complex, feedback-driven dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。