用多尺度注意力网络提升机器人对复杂扣件的精准定位能力
SMR-Net:Robot Snap Detection Based on Multi-Scale Features and Self-Attention Network
- 融合多尺度特征与自注意力机制,增强关键信号抑制噪声
- 在两种扣件数据集上IoU提升超5%,mAP提升超1.5%
- 适合高精度自动化装配场景,尤其透明/低对比度扣件检测
在机器人自动装配中,扣件装配的精度与效率直接影响整体生产质量。作为核心前提,扣件检测与定位直接决定后续装配成败。传统视觉方法在处理复杂场景(如透明或低对比度扣件)时鲁棒性差、定位误差大,难以满足高精度装配需求。为此,本文设计专用传感器并提出SMR-Net——一种基于自注意力的多尺度目标检测算法,协同提升检测与定位性能。SMR-Net采用注意力增强的多尺度特征融合架构:原始传感器数据通过嵌入注意力的特征提取器编码,强化关键扣件特征并抑制噪声;三个多尺度特征图并行经标准与空洞卷积处理以统一维度并保留分辨率;自适应重加权网络动态分配融合特征权重,生成融合细节与全局语义的精细表征。在Type A和Type B扣件数据集上的实验表明,SMR-Net显著优于传统Faster R-CNN:IoU分别提升6.52%和5.8%,mAP分别提高2.8%和1.5%。充分验证了该方法在复杂扣件检测与定位任务中的优越性。
原文摘要 · Abstract (English)
In robot automated assembly, snap assembly precision and efficiency directly determine overall production quality. As a core prerequisite, snap detection and localization critically affect subsequent assembly success. Traditional visual methods suffer from poor robustness and large localization errors when handling complex scenarios (e.g., transparent or low-contrast snaps), failing to meet high-precision assembly demands. To address this, this paper designs a dedicated sensor and proposes SMR-Net, an self-attention-based multi-scale object detection algorithm, to synergistically enhance detection and localization performance. SMR-Net adopts an attention-enhanced multi-scale feature fusion architecture: raw sensor data is encoded via an attention-embedded feature extractor to strengthen key snap features and suppress noise; three multi-scale feature maps are processed in parallel with standard and dilated convolution for dimension unification while preserving resolution; an adaptive reweighting network dynamically assigns weights to fused features, generating fine representations integrating details and global semantics. Experimental results on Type A and Type B snap datasets show SMR-Net outperforms traditional Faster R-CNN significantly: Intersection over Union (IoU) improves by 6.52% and 5.8%, and mean Average Precision (mAP) increases by 2.8% and 1.5% respectively. This fully demonstrates the method's superiority in complex snap detection and localization tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。