arXiv:2608.04999eess.SYcs.AI2026-08被引 1

用大模型引导的多目标强化学习,让模拟电路设计更快更准。

ORACLE: A Multi-Objective Reinforcement Learning-Based Analog Circuit Design Optimizer with Large Language Models-Guided Exploration

  • 采用向量奖励和偏好条件化,实现真正多目标优化。
  • 测试中运行时间缩短20.4至104.4倍,达标率99.9%。
  • 单模型适配多种权衡需求,适合芯片设计与自动化研究者。

基于强化学习(RL)的模拟电路设计自动化已成为降低人工成本的有前景方法。然而,现有大多数方法聚焦于单目标优化,即使针对多目标(MO)问题的方法也常将多个设计规格简化为单一标量奖励,这限制了对真实帕累托前沿的捕捉能力,导致设计性能不佳。此外,每当目标规格变化时需重新训练模型仍是关键瓶颈。为此,我们提出ORACLE,一个开源的基于强化学习的多目标模拟电路设计优化框架,以向量值学习和偏好感知条件化替代标量奖励优化。ORACLE使用偏好向量指定多个目标的相对权重,使单个训练好的模型可在无需重训的情况下生成多种权衡配置下的设计。我们还提出了两种偏好引导策略:归一化权重引导和余弦对齐引导,以提升收敛性。此外,引入大语言模型(LLM)引导的动作选择机制,过滤可能导致次优设计或增加运行时间的动作。在多个电路拓扑结构上共2,000个测试案例中,ORACLE相比最先进方法运行时间减少20.4x–104.4x,满足99.9%的目标规格,并在输出规格的指标上取得5.1x–318.6x的提升。

原文摘要 · Abstract (English)

Analog circuit design automation using reinforcement learning (RL) has emerged as a promising approach for reducing manual effort. However, many existing RL-based methods focus on single-objective optimization. Even methods designed for multi-objective (MO) problems often reduce multiple design specifications to a single scalar reward. This simplification limits the ability to capture the true Pareto trade-off among competing objectives and often leads to suboptimal designs. Moreover, requiring the model to be retrained from scratch whenever the desired MO specifications change remains a key limitation. To address these challenges, we present ORACLE, an open-source RL-based framework for MO analog circuit design optimization that replaces scalar reward optimization with vector-valued learning and preference-aware conditioning. ORACLE represents a true MO analog circuit design optimizer that uses a preference vector to specify the relative weights of multiple objectives, enabling a single trained model to generate designs across diverse trade-off settings without retraining. We further propose two preference-guidance strategies, namely normalized-weight guidance and cosine-aligned guidance, to improve convergence. In addition, we incorporate a large language model (LLM)-guided action selection mechanism to filter actions that are likely to lead to suboptimal designs or increased runtime. Our results show that, on multiple circuit topologies with 2,000 test cases, ORACLE reduces runtime by 20.4x - 104.4x compared to state-of-the-art approaches. It also meets 99.9% of the 2,000 target specifications, and achieves 5.1x - 318.6x better figure of merit in the resulting output specs.

模拟电路多目标优化强化学习大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。