用双层模型优化激光沉积扫描顺序,提升成形质量。
Reinforcement Learning with a Bilevel World-Model Architecture for Scan-Order Optimisation in Laser Directed Energy Deposition

- 构建双层架构:教师模型指导策略学习,生成合法扫描路径。
- PPO策略在小轨迹数下表现稳定,大序列时可靠性下降。
- 物理约束优先,实现可解释的全局分散与局部聚合扫描模式。
激光定向能量沉积(LDED)中的扫描顺序设计是一个延迟、路径依赖的热力耦合决策问题,因序列质量需在完整沉积与冷却后才能评估。本文将该问题建模为有限时域、排列约束的强化学习任务,提出基于有限元教师标签的双层AI优化流程。通过代理辅助的教师引导优化循环学习Abaqus标注的响应景观,为策略训练提供可处理的终态奖励环境。冻结的可屏蔽近端策略优化(MaskablePPO)策略生成合法扫描顺序,并由Abaqus热力仿真独立验证。结果表明,策略生成价值受规模(N)限制,未超越成熟代理优化器;最优序列由教师引导的代理循环获得,而PPO自主探索原响应景观中具有竞争力的区域,在较小轨迹数下秩集中更优,长时域时出现明确可靠性边界。教师标注景观支持物理门控的词典式奖励层级:变形可接受性为首要约束,塑性应变为安全过滤器,残余应力改善仅在可接受区域内有条件追求。验证序列揭示出可解释的尺度分离排序倾向,结合全局空间分散与局部结构分组。该工作为从固定扫描规则向有限元教师验证的策略生成提供路径,同时保留有限元验证作为最终物理校验关卡。
原文摘要 · Abstract (English)
Scan-order design in laser directed energy deposition (LDED) is a delayed, path-dependent thermo-mechanical decision problem, because sequence quality becomes observable only after the complete deposition and cooling cycle. This work formulates LDED scan-order optimisation as a finite-horizon, permutation-constrained reinforcement-learning problem and develops a bilevel finite-element-teacher-labelled AI workflow. A surrogate-assisted teacher-guided optimisation loop learns the Abaqus-labelled response landscape and provides a tractable terminal-reward environment for policy training. A frozen Maskable Proximal Policy Optimization (MaskablePPO) policy is then used to generate legal scan-order candidates, which are independently validated through Abaqus thermo-mechanical simulations. The results show bounded, N-dependent policy-generation value rather than record-level dominance over the mature surrogate-assisted optimiser. The strongest scan orders are obtained by the teacher-guided surrogate loop, whereas PPO autonomously reaches competitive regions of the native response landscape, with stronger rank concentration at smaller track counts and a clear reliability boundary at longer horizons. The teacher-labelled landscape further supports a physically gated lexicographic reward hierarchy in which warpage admissibility is the primary constraint, plastic strain acts as a safety filter and residual-stress-related improvement is pursued conditionally within the admissible region. Validated sequences also reveal an interpretable scale-separated ordering tendency that combines global spatial dispersion with local structured grouping. This workflow provides a route from fixed scan-rule selection toward finite-element-teacher-validated policy generation, while preserving independent finite-element validation as the final physical gate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。