arXiv:2603.19583cs.SEcs.AI2026-03被引 2

让AI在嵌入式系统开发中像专家一样精准操作硬件。

Skilled AI Agents for Embedded and IoT Systems Development

  • 用技能模块化封装硬件操作经验,提升AI编程可靠性。
  • 实测378次硬件运行,专家技能使成功率接近100%。
  • 适合做物联网和嵌入式开发的AI研究者与工程师参考。

大型语言模型(LLMs)和智能体系统在自动化软件开发中展现出潜力,但将其应用于软硬件紧密耦合的硬件在回路(HIL)嵌入式与物联网(IoT)系统仍面临挑战,因为编译通过的代码在真实设备上可能因时序约束、外设初始化要求或硬件特异性行为而失败。为此,我们提出一种基于技能的智能体框架用于HIL嵌入式开发,并构建了IoT-SkillsBench基准,系统评估AI智能体在真实嵌入式环境中的表现。该基准覆盖三个典型嵌入式平台、23种外设及42项任务,分三个难度等级,每项任务在三种配置下评估(无技能、LLM生成技能、人类专家技能),并通过真实硬件执行验证。在378次硬件验证实验中,结构化的专家技能显著提升成功率,实现跨平台近满分表现。

原文摘要 · Abstract (English)

Large language models (LLMs) and agentic systems have shown promise for automated software development, but applying them to hardware-in-the-loop (HIL) embedded and Internet-of-Things (IoT) systems remains challenging due to the tight coupling between software logic and physical hardware behavior. Code that compiles successfully may still fail when deployed on real devices because of timing constraints, peripheral initialization requirements, or hardware-specific behaviors. To address this challenge, we introduce a skills-based agentic framework for HIL embedded development together with IoT-SkillsBench, a benchmark designed to systematically evaluate AI agents in real embedded programming environments. IoT-SkillsBench spans three representative embedded platforms, 23 peripherals, and 42 tasks across three difficulty levels, where each task is evaluated under three agent configurations (no-skills, LLM-generated skills, and human-expert skills) and validated through real hardware execution. Across 378 hardware validated experiments, we show that concise human-expert skills with structured expert knowledge enable near-perfect success rates across platforms.

嵌入式AI开发物联网智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。