arXiv:2607.17950cs.RO2026-07

统一网络同时检测行李车位置、关键点和朝向,提速降耗。

UMCP: A Unified Multi-Task Collaborative Perception Network for Luggage Trolley Pose Estimation

论文配图:UMCP: A Unified Multi-Task Collaborative Perception Network for Luggage Trolley Pose Estimation
图 1 · 摘自论文原文
  • 用统一网络协同完成检测、关键点与朝向估计
  • 相比现有方法模型更轻,计算量减少37.5%
  • 适合部署在资源受限的机器人上

在机器人自主行李车收集任务中,机器人需在杂乱动态环境中持续定位散落的行李车,要求视觉系统兼具高精度与实时性。然而,现有行李车视觉感知方法多依赖级联的多模型推理,导致推理延迟增加和部署成本上升。为此,本文提出统一多任务协作感知网络(UMCP),可同步完成行李车检测、关键点检测与朝向估计。基于YOLOv12架构,将关键点特征与朝向特征融合,并输入朝向特征增强模块(OFEM),提升朝向估计精度。此外,采用圆形概率分布建模与Kullback-Leibler(KL)散度损失进一步优化朝向估计。实验表明,该方法在保持竞争性整体精度的同时,显著降低模型复杂度与计算开销,相较现有方法计算量减少37.5%。相关项目主页见:https://sites.google.com/view/robot-umcp。

原文摘要 · Abstract (English)

In robotic autonomous luggage trolley collection, robots must continuously localize scattered luggage trolleys in cluttered and dynamic environments. This requires the vision system to achieve both high accuracy and real-time performance. However, existing visual perception approaches for luggage trolleys often rely on cascaded multi-model inference, leading to increased inference latency and high deployment costs. To address these limitations, this article presents a unified multi-task collaborative perception network (UMCP) that simultaneously performs luggage trolley detection, keypoint detection and orientation estimation. Based on the YOLOv12 architecture, keypoint features are fused with orientation features and then fed into an orientation feature enhancement module (OFEM), thereby improving orientation estimation accuracy. In addition, circular probability distribution modeling with a Kullback-Leibler (KL) divergence loss is adopted to enhance orientation estimation accuracy further. Experimental results demonstrate that the proposed method achieves competitive overall accuracy while substantially reducing model complexity and computational cost compared with existing methods. A website about this work is available at https://sites.google.com/view/robot-umcp.

多任务学习姿态估计机器人感知轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。