arXiv:2411.16019cs.LG2024-11

用Mamba+调度策略,让强化学习更快优化多种模拟电路。

M3: Mamba-assisted Multi-Circuit Optimization via MBRL with Effective Scheduling

  • 用Mamba替代Transformer,统一处理多类电路参数与目标
  • 通过动态调整训练参数,样本效率提升显著
  • 适合需要快速适配新电路的芯片设计人员

基于模型的强化学习(MBRL)在模拟电路优化中展现出提升采样效率和跨拓扑泛化能力的潜力。然而,仍面临计算开销高、每类电路需定制模型等挑战。为此,我们提出M3,一种结合Mamba架构与有效调度策略的新型MBRL方法。Mamba作为Transformer的高效替代,可统一处理不同参数与目标规格的多电路优化。有效调度策略通过动态调节关键训练参数,显著提升采样效率。据我们所知,M3是首个同时利用Mamba架构与有效调度实现多电路优化的方案,在采样效率上优于现有RL方法。

原文摘要 · Abstract (English)

Recent advancements in reinforcement learning (RL) for analog circuit optimization have demonstrated significant potential for improving sample efficiency and generalization across diverse circuit topologies and target specifications. However, there are challenges such as high computational overhead, the need for bespoke models for each circuit. To address them, we propose M3, a novel Model-based RL (MBRL) method employing the Mamba architecture and effective scheduling. The Mamba architecture, known as a strong alternative to the transformer architecture, enables multi-circuit optimization with distinct parameters and target specifications. The effective scheduling strategy enhances sample efficiency by adjusting crucial MBRL training parameters. To the best of our knowledge, M3 is the first method for multi-circuit optimization by leveraging both the Mamba architecture and a MBRL with effective scheduling. As a result, it significantly improves sample efficiency compared to existing RL methods.

强化学习电路优化Mamba采样效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。