arXiv:2608.14579cs.AI2026-08

用大模型迭代优化逻辑综合,自动纠错提升成功率。

SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization

  • 三模型协同:规划、推理、分析分工,结合强化学习交互优化
  • 在500K门电路上达86.3%成功率,比人工流程提升12.4%性能
  • 自纠错机制实时检测问题并恢复,适合复杂电路自动化设计

逻辑综合优化因搜索空间指数级增长、奖励信号稀疏及结构多样而面临挑战。传统专家设计流程适应性差,强化学习方法常存在样本效率低、可解释性弱的问题。本文提出SKILL,一种自纠正知识引导的迭代式大语言模型智能体,融合多代理大模型推理与基于强化学习的环境交互,实现自动化综合优化。SKILL协调三个专用LLM:GPT-4o负责战略规划,Claude Sonnet 4进行详细推理,Gemini 2.5 Pro高效分析,并搭配基于PPO的强化学习代理直接与综合工具交互学习可行策略。创新的自纠正模块监控环境反馈(PDA指标),识别次优行为并触发大模型引导的恢复策略。在IWLS、OpenCores和EPFL基准上的评估显示,SKILL相比专家流程实现12.4%的PDA提升,在500K门规模逻辑系统上达到86.3%的成功率。

原文摘要 · Abstract (English)

Logic synthesis optimization poses significant challenges due to exponentially growing search spaces, sparse reward signals, and diverse logic structures. Traditional expert-designed flows lack adaptability, while reinforcement learning (RL) methods often suffer from low sample efficiency and limited interpretability. We introduce SKILL, a Self-correcting Knowledge-guided Iterative Large Language Model Agent that unifies multi-agent LLM reasoning and RL-based environment interaction for automated synthesis optimization. SKILL coordinates three specialized LLMs: GPT-4o for strategic planning, Claude Sonnet 4 for detailed reasoning, and Gemini 2.5 Pro for efficient analysis with a PPO-based RL agent that learns actionable policies through direct interaction with synthesis tools. A novel self-correcting module monitors environment feedback (PDA metrics), detects suboptimal behaviors, and invokes LLM-guided recovery strategies. Evaluations on IWLS, OpenCores, and EPFL benchmarks show SKILL achieves a 12.4 % PDA improvement over expert flows and 86.3% success rate on logic systems up to 500K gates.

逻辑优化大模型强化学习自动化设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。