BlabberSeg让无人机实时识别空中任意物体,速度提升超9倍
BlabberSeg: Real-Time Embedded Open-Vocabulary Aerial Segmentation
- 复用提示与模型特征,优化CLIPSeg结构提升效率
- 在Jetson Orin上达16.78帧/秒,速度提升927.41%
- 适合需要低延迟、高适应性的无人机视觉系统
实时空中图像分割对无人飞行器(UAV)环境感知至关重要。本文提出BlabberSeg,一种基于CLIPSeg优化的视觉语言模型,专为机载实时处理航空图像设计。通过复用提示与模型特征,大幅降低计算开销,实现真正的开放词汇空中分割。我们在安全着陆场景中验证了BlabberSeg,采用视觉伺服与开放词汇分割的DOVESEI框架。相较于原始CLIPSeg(1.81 Hz),BlabberSeg在NVIDIA Jetson Orin AGX(64GB)上实现16.78 Hz的推理速度,提速927.41%,准确率损失仅2.1%(正确分割区域占比)。代码已开源。
原文摘要 · Abstract (English)
Real-time aerial image segmentation plays an important role in the environmental perception of Uncrewed Aerial Vehicles (UAVs). We introduce BlabberSeg, an optimized Vision-Language Model built on CLIPSeg for on-board, real-time processing of aerial images by UAVs. BlabberSeg improves the efficiency of CLIPSeg by reusing prompt and model features, reducing computational overhead while achieving real-time open-vocabulary aerial segmentation. We validated BlabberSeg in a safe landing scenario using the Dynamic Open-Vocabulary Enhanced SafE-Landing with Intelligence (DOVESEI) framework, which uses visual servoing and open-vocabulary segmentation. BlabberSeg reduces computational costs significantly, with a speed increase of 927.41% (16.78 Hz) on a NVIDIA Jetson Orin AGX (64GB) compared with the original CLIPSeg (1.81Hz), achieving real-time aerial segmentation with negligible loss in accuracy (2.1% as the ratio of the correctly segmented area with respect to CLIPSeg). BlabberSeg's source code is open and available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。