用强化学习直接求解原子非局部热平衡辐射传输问题,无需近似算子。
Policy-Based Radiative Transfer: Solving the $2$-Level Atom Non-LTE Problem using Soft Actor-Critic Reinforcement Learning
- 将辐射传输建模为强化学习控制任务,让智能体自主学习深度相关的源函数。
- 在不依赖近似算子和标注数据的情况下,实现统计平衡方程的自洽求解。
- 适用于复杂大气场景,可调节奖励权重以优化求解效率,适合太阳物理研究者。
我们提出一种新型强化学习方法,将经典的二能级原子非局部热平衡辐射传输问题转化为控制任务:一个强化学习智能体通过与辐射传输引擎的奖励反馈交互,学习深度相关的源函数 $S(τ)$,使其自洽满足统计平衡方程(SE)。该方法完全无需构造常用的加速迭代方案中的近似λ算子($Λ^*$),也无需预先构建大规模标注数据集来获取监督信号,同时避免了对复杂辐射传输求解器进行反向传播。实验表明,仅使用贪婪训练的前馈神经网络无法解决统计平衡问题,可能源于目标动态变化的特性。所提出的 $Λ^*- ext{Free}$ 方法在存在增强速度场、多维几何或复杂微观物理等复杂场景中具有潜在优势,且可通过调整折扣因子激励智能体寻找更高效的策略。若能证明其泛化能力,该框架有望成为实现统计平衡的一种替代或加速形式。据我们所知,这是首次将强化学习直接应用于太阳物理中求解基本物理约束的研究。
原文摘要 · Abstract (English)
We present a novel reinforcement learning (RL) approach for solving the classical 2-level atom non-LTE radiative transfer problem by framing it as a control task in which an RL agent learns a depth-dependent source function $S(τ)$ that self-consistently satisfies the equation of statistical equilibrium (SE). The agent's policy is optimized entirely via reward-based interactions with a radiative transfer engine, without explicit knowledge of the ground truth. This method bypasses the need for constructing approximate lambda operators ($Λ^*$) common in accelerated iterative schemes. Additionally, it requires no extensive precomputed labeled datasets to extract a supervisory signal, and avoids backpropagating gradients through the complex RT solver itself. Finally, we show through experiment that a simple feedforward neural network trained greedily cannot solve for SE, possibly due to the moving target nature of the problem. Our $Λ^*-\text{Free}$ method offers potential advantages for complex scenarios (e.g., atmospheres with enhanced velocity fields, multi-dimensional geometries, or complex microphysics) where $Λ^*$ construction or solver differentiability is challenging. Additionally, the agent can be incentivized to find more efficient policies by manipulating the discount factor, leading to a reprioritization of immediate rewards. If demonstrated to generalize past its training data, this RL framework could serve as an alternative or accelerated formalism to achieve SE. To the best of our knowledge, this study represents the first application of reinforcement learning in solar physics that directly solves for a fundamental physical constraint.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。