arXiv:2506.03607cs.CV2025-06被引 1

让图像描述模型在边缘设备上快速运行,兼顾速度与准确度。

Analyzing Transformer Models and Knowledge Distillation Approaches for Image Captioning on Edge AI

  • 采用轻量Transformer模型结合知识蒸馏技术
  • 在资源受限设备上实现推理加速,性能损失小于5%
  • 适合工业机器人、智能巡检等实时场景

边缘计算将处理能力下沉至网络边缘,推动物联网应用中的实时AI决策。在工业自动化如机器人和坚固型边缘AI中,实时感知与智能对自主运行至关重要。部署基于Transformer的图像描述模型可提升机器感知能力,增强自主机器人对场景的理解,并辅助工业检测。然而,这些边缘或物联网设备通常在计算资源上受限,以保证物理敏捷性,同时又有严格的响应时间要求。传统深度学习模型往往过于庞大且计算开销高,难以适应此类设备。本研究评估了适用于边缘设备的高效Transformer模型,并应用知识蒸馏技术,证明在资源受限设备上通过该方法可实现推理加速,同时保持模型性能。实验表明,经过优化的模型在边缘设备上仍能维持较高生成质量,满足实际部署需求。

原文摘要 · Abstract (English)

Edge computing decentralizes processing power to network edge, enabling real-time AI-driven decision-making in IoT applications. In industrial automation such as robotics and rugged edge AI, real-time perception and intelligence are critical for autonomous operations. Deploying transformer-based image captioning models at the edge can enhance machine perception, improve scene understanding for autonomous robots, and aid in industrial inspection. However, these edge or IoT devices are often constrained in computational resources for physical agility, yet they have strict response time requirements. Traditional deep learning models can be too large and computationally demanding for these devices. In this research, we present findings of transformer-based models for image captioning that operate effectively on edge devices. By evaluating resource-effective transformer models and applying knowledge distillation techniques, we demonstrate inference can be accelerated on resource-constrained devices while maintaining model performance using these techniques.

边缘计算图像描述知识蒸馏Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。