arXiv:2503.18229cs.LGcs.AI2025-03被引 1

用多个低精度模型动态辅助高精度模型,降低工程优化中的学习方差。

Adaptive Multi-Fidelity Reinforcement Learning for Variance Reduction in Engineering Design Optimization

  • 通过动态融合多个非层级低精度模型提升学习效率
  • 在八旋翼设计中使策略学习方差显著下降,收敛更快
  • 无需人工调参,减少计算负担,适合复杂工程优化

多保真度强化学习框架通过整合不同精度与成本的分析模型,高效利用计算资源。现有方法多依赖模型层级结构,但在设计空间中各模型误差分布不均时,易加剧策略学习方差。本文提出一种新型自适应多保真度RL框架,动态融合多个异构、非层级的低保真度模型与一个高保真度模型,以高效学习高保真度策略。具体而言,低保真度策略及其经验数据根据其与高保真度策略的一致性被自适应地用于聚焦学习。在八旋翼设计优化问题中,使用两个低保真度模型与一个高保真度模拟器进行验证,结果表明该方法显著降低策略学习方差,提升收敛速度与解的质量,优于传统层级式多保真度方法。此外,框架无需手动设定模型使用调度,避免了额外计算开销。该方法为多保真度强化学习提供了一种有效的方差抑制策略,同时减轻了人工调度带来的计算与操作负担。

原文摘要 · Abstract (English)

Multi-fidelity Reinforcement Learning (RL) frameworks efficiently utilize computational resources by integrating analysis models of varying accuracy and costs. The prevailing methodologies, characterized by transfer learning, human-inspired strategies, control variate techniques, and adaptive sampling, predominantly depend on a structured hierarchy of models. However, this reliance on a model hierarchy can exacerbate variance in policy learning when the underlying models exhibit heterogeneous error distributions across the design space. To address this challenge, this work proposes a novel adaptive multi-fidelity RL framework, in which multiple heterogeneous, non-hierarchical low-fidelity models are dynamically leveraged alongside a high-fidelity model to efficiently learn a high-fidelity policy. Specifically, low-fidelity policies and their experience data are adaptively used for efficient targeted learning, guided by their alignment with the high-fidelity policy. The effectiveness of the approach is demonstrated in an octocopter design optimization problem, utilizing two low-fidelity models alongside a high-fidelity simulator. The results demonstrate that the proposed approach substantially reduces variance in policy learning, leading to improved convergence and consistent high-quality solutions relative to traditional hierarchical multi-fidelity RL methods. Moreover, the framework eliminates the need for manually tuning model usage schedules, which can otherwise introduce significant computational overhead. This positions the framework as an effective variance-reduction strategy for multi-fidelity RL, while also mitigating the computational and operational burden of manual fidelity scheduling.

强化学习多保真度工程优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。