arXiv:2605.00963cs.ROcs.AI2026-05

拆解人机交互中感知、语言、控制三模块,找出影响成功率和执行时间的关键因素。

Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task

论文配图:Ablation Study of Multimodal Perception, Language Grounding, and Control for Human-Robot Interaction in an Object Detection and Grasping Task
图 1 · 摘自论文原文
  • 通过受控实验分离语言模型、视觉感知与运动控制器的影响。
  • 对比3种语言模型、5种感知配置和3种控制器,找到最优组合。
  • 揭示各模块对任务性能的贡献,指导未来系统优化方向。

本文在先前多模态人机交互系统基础上,引入受控消融实验,分析三个对端到端性能影响最显著的模块:用于动作提取的大语言模型、用于视觉定位的感知系统,以及用于运动执行的控制器。目标并非重构整个流程,而是通过统一实验协议,隔离各组件的贡献,并最终评估最佳组合的端到端表现。因此,我们比较了三种语言模型、五种感知配置和三种控制器,随后对最优候选方案进行二次因子分析。该分析旨在明确哪些选择主要影响执行时间,哪些主要影响成功概率,并指出未来系统改进中最可能带来工程收益的环节。

原文摘要 · Abstract (English)

This manuscript extends our previous multimodal human-robot interaction system by introducing a controlled ablation study of the three modules that most strongly influence end-to-end performance: the large language model used for action extraction, the perception system used for visual grounding, and the controller used for motion execution. The goal is not to redesign the full pipeline, but to isolate the contribution of each component under a common experimental protocol and then evaluate the best combinations end-to-end. We therefore compare three language models, five perception configurations, and three controllers, followed by a second-stage factorial study over the best candidates. The resulting analysis is intended to clarify which choices primarily affect execution time, which primarily affect success rate, and where the largest engineering gains are likely to come from in future revisions of the system.

人机交互多模态消融实验机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。