仿生注意力提升小目标定位精度,实测准确率提升超4%。
CA-YOLO: Cross Attention Empowered YOLO for Biomimetic Localization
- 引入仿生注意力模块与小目标检测头,模拟动物视觉聚焦机制。
- 在COCO和VisDrone数据集上准确率分别提升3.94%和4.90%。
- 适合需高精度小目标定位的智能监控、机器人导航场景。
在复杂环境中实现精准高效的目标定位至关重要。现有系统常面临精度不足与小目标识别能力弱的问题。本文提出基于CA-YOLO的仿生稳定定位系统,通过在YOLO主干网络中嵌入仿生模块,增强目标定位精度与小目标识别能力。该算法模仿动物视觉聚焦机制,引入小目标检测头及特征融合注意力机制(CFAM)。同时,借鉴人类前庭-眼反射(VOR),设计仿生云台跟踪控制策略,包含中心定位、稳定性优化、自适应控制系数调节与智能重捕功能。实验结果表明,CA-YOLO在标准数据集COCO和VisDrone上平均精度分别提升3.94%和4.90%。时敏目标定位实验进一步验证了系统的有效性与实用性。
原文摘要 · Abstract (English)
In modern complex environments, achieving accurate and efficient target localization is essential in numerous fields. However, existing systems often face limitations in both accuracy and the ability to recognize small targets. In this study, we propose a bionic stabilized localization system based on CA-YOLO, designed to enhance both target localization accuracy and small target recognition capabilities. Acting as the "brain" of the system, the target detection algorithm emulates the visual focusing mechanism of animals by integrating bionic modules into the YOLO backbone network. These modules include the introduction of a small target detection head and the development of a Characteristic Fusion Attention Mechanism (CFAM). Furthermore, drawing inspiration from the human Vestibulo-Ocular Reflex (VOR), a bionic pan-tilt tracking control strategy is developed, which incorporates central positioning, stability optimization, adaptive control coefficient adjustment, and an intelligent recapture function. The experimental results show that CA-YOLO outperforms the original model on standard datasets (COCO and VisDrone), with average accuracy metrics improved by 3.94%and 4.90%, respectively.Further time-sensitive target localization experiments validate the effectiveness and practicality of this bionic stabilized localization system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。