arXiv:2411.10841cs.LGcs.AI2024-11被引 1

用自适应策略融合非分层多精度模型,提升工程设计效率。

Adaptive Learning of Design Strategies over Non-Hierarchical Multi-Fidelity Models via Policy Alignment

  • 通过策略对齐动态选择低精度模型辅助学习
  • 在测试优化与八旋翼设计中收敛更快,性能更优
  • 无需预设模型调度,适合复杂设计场景

多精度强化学习框架通过使用不同精度和计算成本的分析模型,显著提升工程设计效率。现有方法大多依赖模型层级结构,但忽略了模型在设计空间中误差分布的异质性。本文提出ALPHA(Adaptively Learned Policy with Heterogeneous Analyses),一种新型多精度强化学习框架,能够自适应地利用任意非层级、异构的低精度模型与一个高精度模型,共同学习高性能策略。具体而言,低精度策略及其经验数据被动态用于高效目标学习,其使用由与高精度策略的对齐程度决定。在解析测试优化与八旋翼设计问题中的实验表明,ALPHA能随时间和设计空间自适应调用模型,无需如层级框架般预设模型调度。此外,自适应代理能发现更直接的高性能解路径,展现出优于层级代理的收敛性能。

原文摘要 · Abstract (English)

Multi-fidelity Reinforcement Learning (RL) frameworks significantly enhance the efficiency of engineering design by leveraging analysis models with varying levels of accuracy and computational costs. The prevailing methodologies, characterized by transfer learning, human-inspired strategies, control variate techniques, and adaptive sampling, predominantly depend on a structured hierarchy of models. However, this reliance on a model hierarchy overlooks the heterogeneous error distributions of models across the design space, extending beyond mere fidelity levels. This work proposes ALPHA (Adaptively Learned Policy with Heterogeneous Analyses), a novel multi-fidelity RL framework to efficiently learn a high-fidelity policy by adaptively leveraging an arbitrary set of non-hierarchical, heterogeneous, low-fidelity models alongside a high-fidelity model. Specifically, low-fidelity policies and their experience data are dynamically used for efficient targeted learning, guided by their alignment with the high-fidelity policy. The effectiveness of ALPHA is demonstrated in analytical test optimization and octocopter design problems, utilizing two low-fidelity models alongside a high-fidelity one. The results highlight ALPHA's adaptive capability to dynamically utilize models across time and design space, eliminating the need for scheduling models as required in a hierarchical framework. Furthermore, the adaptive agents find more direct paths to high-performance solutions, showing superior convergence behavior compared to hierarchical agents.

强化学习多精度建模工程优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。