用云边协同提升机器人在复杂环境中的手势识别与任务执行能力
An Intelligent-Cloud Edge Multimodal Interaction System for Robots

- 云边协同架构,融合改进YOLO检测与大模型代理实现多模态交互
- 小手势或遮挡场景下精度达98.9%,任务成功率最高95%
- 适合资源受限机器人应用,尤其需高精度人机交互的场景
复杂环境中实现可靠的机器人-人类交互,需要精确的手势感知、语义场景理解以及在有限本地算力下的稳定任务规划。本文提出一种云-边多模态交互框架,集成改进的基于YOLO的手势检测器与协同的大语言模型(LLM)和视觉-语言模型(VLM)代理。所提检测器在颈部引入卷积块注意力模块(CBAM),并以Distance-IoU(DIoU)损失替代原始边界框回归目标,提升了小尺寸或部分遮挡手势在复杂背景下的特征区分与定位能力。云端负责手势检测、场景理解、多模态融合与动作规划,而TonyPi机器人本地完成数据采集、通信、动作执行与反馈。在公开手势数据集和自建数据集上的实验表明,YOLO-DC分别达到98.9%和95.0%的精度,[email protected]分别为90.7%和92.7%。系统级评估显示单动作、复合动作及视觉依赖任务的成功率分别为95%、88%和82%。30名参与者评估获得平均满意度3.69/5。结果验证了精细手势检测与多模态代理结合在资源受限机器人交互中的可行性。
原文摘要 · Abstract (English)
Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources. This paper presents a cloud-edge multimodal interaction framework that integrates an enhanced YOLO-based gesture detector with coordinated large language model (LLM) and vision-language model (VLM) agents. The proposed detector, incorporates the Convolutional Block Attention Module (CBAM) into the neck and replaces the baseline bounding-box regression objective with Distance-IoU (DIoU) loss. These modifications improve feature discrimination and localization for small or partially occluded gestures in complex backgrounds. The cloud layer performs gesture detection, scene understanding, multimodal fusion, and action planning, whereas the TonyPi robot locally handles data acquisition, communication, action execution, and feedback. Experiments on a public gesture dataset and a custom dataset show that YOLO-DC achieves precision values of 98.9% and 95.0%, with [email protected] values of 90.7% and 92.7%, respectively. System-level evaluation yields success rates of 95%, 88%, and 82% for single-action, composite-action, and vision-dependent tasks. A 30 participant evaluation yields an overall mean satisfaction score of 3.69 out of 5. These results demonstrate the feasibility of combining refined gesture detection with multimodal agents for resource-constrained robotic interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。