arXiv:2506.03474cs.LGcs.AI2025-06被引 1

CORE用一步强化学习加速芯片设计,兼顾效率与约束满足。

CORE: Constraint-Aware One-Step Reinforcement Learning for Simulation-Guided Neural Network Accelerator Design

  • 基于结构化分布采样设计配置,用图解码器建模依赖关系。
  • 无需价值函数,通过批量奖励对比提升采样效率,显著减少无效设计。
  • 适用于复杂约束的软硬件协同设计,适合芯片架构师和自动化设计研究者。

基于仿真的设计空间探索(DSE)旨在高效优化高维结构化设计,同时应对复杂约束和高昂评估成本。现有方法包括启发式和多步强化学习(RL),因反馈稀疏、延迟以及混合动作空间大,难以平衡采样效率与约束满足。本文提出CORE,一种面向仿真引导的约束感知型一步强化学习方法。在CORE中,策略智能体通过定义设计配置的结构化分布进行采样,利用基于缩放图的解码器建模变量间依赖,并通过奖励塑形惩罚无效设计。策略更新采用替代目标,比较批次内设计的奖励,不需学习价值函数。该无评论家形式使学习更高效,推动选择更高奖励的设计。我们将CORE应用于神经网络加速器的硬件映射联合设计,结果表明其显著提升采样效率,并优于现有最优基线,获得更优的加速器配置。该方法具通用性,可推广至各类离散-连续约束设计问题。

原文摘要 · Abstract (English)

Simulation-based design space exploration (DSE) aims to efficiently optimize high-dimensional structured designs under complex constraints and expensive evaluation costs. Existing approaches, including heuristic and multi-step reinforcement learning (RL) methods, struggle to balance sampling efficiency and constraint satisfaction due to sparse, delayed feedback, and large hybrid action spaces. In this paper, we introduce CORE, a constraint-aware, one-step RL method for simulationguided DSE. In CORE, the policy agent learns to sample design configurations by defining a structured distribution over them, incorporating dependencies via a scaling-graph-based decoder, and by reward shaping to penalize invalid designs based on the feedback obtained from simulation. CORE updates the policy using a surrogate objective that compares the rewards of designs within a sampled batch, without learning a value function. This critic-free formulation enables efficient learning by encouraging the selection of higher-reward designs. We instantiate CORE for hardware-mapping co-design of neural network accelerators, demonstrating that it significantly improves sample efficiency and achieves better accelerator configurations compared to state-of-the-art baselines. Our approach is general and applicable to a broad class of discrete-continuous constrained design problems.

强化学习芯片设计约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。