用真实硬件反馈训练智能体,自动优化嵌入式设备的模型与固件。
Embedded Arena: Iterative Optimization via Hardware Feedback

- 让大模型在真实硬件上反复编译、烧录、测量,实现闭环优化。
- 视觉模型压缩250倍误差低于3.3%,音频模型压缩400倍误特征率<6%。
- 适合需要低功耗、高可靠性的物联网与医疗可穿戴设备研发者。
从野生动物监测站到临床可穿戴设备,嵌入式设备因延迟、通信或隐私限制需本地运行AI推理。在异构微控制器(MCU)上优化模型需同时满足内存、功耗和温度等硬性物理约束,且保持精度,这一多维优化目前依赖专家手动完成。本文探讨大模型智能体是否能借助真实硬件反馈,自主导航复杂多轮优化流程,并提出一个硬件在环的智能体竞技场。智能体通过反复编译、烧录并测量真实硬件上的表现,实现模型与固件的协同优化。前沿模型如Claude Opus 4.7和Gemini 3.1 Pro在无硬件反馈时部署成功率均为0%,而本方法在三轮内首次成功部署,七轮内超越人类专家。该方法使视觉模型压缩250倍、准确率损失<3.3%,音频模型压缩400倍、特征误差率损失<6%,可在商业级MCU上实现太阳能供电的无电池运行。在两个真实系统中验证:鹿类检测摄像头达到96.7%准确率,语音转录可穿戴设备在儿童发育研究中达到8.44%特征误差率。
原文摘要 · Abstract (English)
Embedded devices from wildlife monitoring stations to clinical wearables require local AI inference due to latency, communication, or privacy constraints. Optimizing models for heterogeneous microcontrollers (MCUs) requires simultaneously satisfying hard physical constraints on memory, power, and temperature while preserving accuracy, a multidimensional optimization that is today performed manually by experts. We ask whether an LLM agent can autonomously navigate this complex, multi-turn pipeline guided by real hardware feedback, and introduce a hardware-in-the-loop agent arena in which the agent iteratively refines both model and firmware -- compiling, flashing, and measuring on real hardware -- to enable closed-loop optimization. Frontier models, including Claude Opus 4.7 and Gemini 3.1 Pro, fail entirely without hardware feedback (0% deployment success), whereas our hardware-in-the-loop formulation achieves the first successful deployment within three iterations and can surpass human expert results within seven. This agentic co-optimization achieves 250x compression for vision models with <3.3% accuracy loss and 400x for audio with <6% Feature Error Rate loss, enabling battery-free operation on a commercial MCU via solar harvesting. We demonstrate practical impact in two real-world systems: an elk-detection camera trap (96.7% accuracy) and a phonetic-transcription wearable (8.44% FER) for child development research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。