arXiv:2509.20688cs.ROcs.CV2025-09被引 2

针对机器人硬件设计轻量模型,兼顾精度与推理速度。

RAM-NAS: Resource-aware Multiobjective Neural Architecture Search Method for Robot Vision Tasks

  • 用子网互蒸馏和解耦知识蒸馏优化超网络预训练
  • 在三类机器人边缘设备上训练延迟预测器,实现精准延迟估计
  • 搜索出的模型在图像识别、检测、分割任务中均提速明显

神经架构搜索(NAS)在自动设计轻量级模型方面展现出巨大潜力。然而,传统方法在超网络预训练上表现不足,且忽视实际机器人硬件资源。为此,我们提出RAM-NAS,一种面向机器人视觉任务的资源感知多目标NAS方法,聚焦于提升超网络预训练效果与硬件资源感知能力。引入子网互蒸馏机制,即通过夹心采样规则生成的全部子网之间相互蒸馏;同时采用解耦知识蒸馏(DKD)损失,增强logits蒸馏性能。为加速搜索过程并考虑硬件资源,我们利用三类机器人边缘硬件的数据训练了延迟代理预测器,可在搜索阶段估计硬件推理延迟,支持统一的多目标进化搜索,以平衡模型精度与延迟。所发现的RAM-NAS模型家族在ImageNet上实现76.7%至81.4%的Top-1准确率。此外,所采用的资源感知多目标NAS显著降低了模型在边缘硬件上的推理延迟。我们在下游任务上进行了实验验证方法的可扩展性,所有三种硬件上检测与分割的推理时间均优于基于MobileNetv3的方法。本工作填补了机器人硬件资源感知NAS的空白。

原文摘要 · Abstract (English)

Neural architecture search (NAS) has shown great promise in automatically designing lightweight models. However, conventional approaches are insufficient in training the supernet and pay little attention to actual robot hardware resources. To meet such challenges, we propose RAM-NAS, a resource-aware multi-objective NAS method that focuses on improving the supernet pretrain and resource-awareness on robot hardware devices. We introduce the concept of subnets mutual distillation, which refers to mutually distilling all subnets sampled by the sandwich rule. Additionally, we utilize the Decoupled Knowledge Distillation (DKD) loss to enhance logits distillation performance. To expedite the search process with consideration for hardware resources, we used data from three types of robotic edge hardware to train Latency Surrogate predictors. These predictors facilitated the estimation of hardware inference latency during the search phase, enabling a unified multi-objective evolutionary search to balance model accuracy and latency trade-offs. Our discovered model family, RAM-NAS models, can achieve top-1 accuracy ranging from 76.7% to 81.4% on ImageNet. In addition, the resource-aware multi-objective NAS we employ significantly reduces the model's inference latency on edge hardware for robots. We conducted experiments on downstream tasks to verify the scalability of our methods. The inference time for detection and segmentation is reduced on all three hardware types compared to MobileNetv3-based methods. Our work fills the gap in NAS for robot hardware resource-aware.

神经架构搜索机器人视觉边缘计算多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。