用大模型指导神经网络架构搜索,适配特殊硬件并提升能效与鲁棒性。
LLM-Guided Neural Architecture Search for Robust Co-Design of Physical Neural Networks

- 引入大模型作为进化算子,实现软硬协同设计
- 在光学MZI硬件上发现更多样且鲁棒的网络结构
- 适合关注新型计算平台的系统设计研究人员
将神经网络部署于非常规硬件需同时优化任务准确率与平台特定约束,如能耗、物理非理想性和数值精度。现有神经架构搜索(NAS)方法通常针对单一硬件家族,限制了跨平台比较与泛化能力。我们提出无硬件依赖的LLM引导NAS框架UH-NAS,通过语言模型作为进化算子,联合优化准确率与推理能耗。通过将硬件设为可替换后端,集成各平台能耗模型、物理约束和非理想性模拟器,UH-NAS可在不修改搜索算法的前提下,实现不同后端间的公平系统级对比。在光学MZI硬件上的测试表明,相比传统基线,UH-NAS发现的架构更具多样性与鲁棒性,且优于现有LLM-to-NAS方法。额外消融实验验证了在非理想条件下架构鲁棒性及系统提示的关键作用,凸显软硬协同设计对新兴计算平台的重要性。
原文摘要 · Abstract (English)
Deploying neural networks on unconventional hardware demands architectures that co-optimize task accuracy and platform-specific constraints such as energy cost, physical non-idealities, and numerical precision. Existing neural architecture search (NAS) methods are typically tailored to a single hardware family, limiting cross-platform comparison and generalization. We introduce Unconventional Hardware Neural Architecture Search (UH-NAS), a hardware-agnostic, LLM-guided NAS framework that integrates language models as evolutionary operators to co-optimize accuracy and inference energy. By exposing hardware as a swappable backend with per-platform energy models, physical constraints, and non-ideality simulators, UH-NAS enables fair system-level comparisons across various backends without modifying the search algorithm. Tested on optical MZI hardware, UH-NAS discovers more diverse, robust architectures than conventional baselines while outperforming existing LLM-to-NAS approaches. Additional ablations on architecture robustness under non-idealities and the role of system prompts highlight the importance of architecture-hardware co-design for emerging computing platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。