通过模仿专家布局,让强化学习实现顶尖芯片布图效果。
How Can Reinforcement Learning Achieve Expert-level Placement?
- 从专家布局反推布图轨迹,构建隐式奖励模型。
- 仅需一个设计即可训练,且在新场景中表现良好。
- 突破传统奖励设计瓶颈,适合芯片物理设计研究者。
芯片布图是物理设计中的关键步骤。尽管基于强化学习的方法近期涌现,但其训练主要聚焦于线长优化,往往难以达到专家水平。我们发现奖励设计是性能差距的主要原因,因此不依赖复杂的规则形式化,而是直接从专家布局中学习,构建奖励模型。该方法从最终的专家布局反推逐步的布图轨迹,利用这些轨迹作为示范或偏好数据,训练模型以捕捉专家结果中的隐式奖励。实验表明,该框架可高效从单一设计中学习,并在未见案例中实现良好泛化。
原文摘要 · Abstract (English)
Chip placement is a critical step in physical design. While reinforcement learning (RL)-based methods have recently emerged, their training primarily focuses on wirelength optimization, and therefore often fail to achieve expert-quality layouts. We identify the reward design as the primary cause for the performance gap with experts, and instead of formalizing intricate processes, we circumvent this by directly learning from expert layouts to derive a reward model. Our approach starts from the final expert layouts to infer step-by-step expert trajectories. Using these trajectories as demonstrations or preferences, we train a model that captures the latent implicit rewards in expert results. Experiments show that our framework can efficiently learn from even a single design and generalize well to unseen cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。