arXiv:2607.00530cs.ROcs.AI2026-07

技术提升15%成功率,用户感知明显更优。

From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping

论文配图:From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping
图 1 · 摘自论文原文
  • 用新模型替换感知与语言模块,提升任务成功率至90%
  • 70.8%用户偏好改进系统,感知速度、可靠性和流畅性均显著提升
  • 验证了技术指标提升可被用户直接感知,强调用户研究的重要性

人机交互系统的技术性能提升,并不自动转化为用户在实际互动中可察觉的差异。本文研究了端到端任务成功率从75%提升至90%(提高15个百分点)是否足以引发用户感知上的显著变化。基线系统采用Whisper进行语音识别,Florence-2实现开放词汇物体检测,LLaMA 3.1提取动作指令,通过区间型二型模糊逻辑控制器执行运动。改进配置则替换为Grounding DINO + SAM进行感知,Qwen 3.5 9B替代语言模块,其余保持不变。对24名参与者开展自身对照实验,在相同桌面上完成抓取任务后,使用7点量表评估感知速度、可靠性及整体能力与流畅性。结果显示,17名参与者(70.83%)偏好改进系统(精确二项检验,p = 0.043, h = 0.43),经霍尔姆校正后,三项感知维度评分均显著更高,效应量为大到极大型(p < 0.001)。结果表明,该技术改进在真实交互中可被用户察觉,强调评估机器人操作流程时应结合用户中心证据补充基准测试。

原文摘要 · Abstract (English)

Improvements in the technical performance of human--robot interaction (HRI) systems do not automatically translate into differences that human users can detect during live interaction. This paper investigates whether a 15 percentage point gain in end-to-end task success (from 75% in a multimodal baseline system to 90% in an improved configuration identified through a prior ablation study) is sufficient to produce consistent and measurable differences in user perception. The baseline system combines Whisper for speech recognition, Florence-2 for open-vocabulary object detection, LLaMA 3.1 for action extraction, and an interval Type-2 fuzzy logic controller for motion execution. The improved configuration replaces the perception and language modules with Grounding DINO + SAM and Qwen 3.5 9B, respectively, while retaining the same controller. A within-subject user study with 24 participants compared both systems on the same tabletop object-grasping task. After interacting with each configuration, participants rated perceived speed, reliability, and overall competence and fluency on a 7-point Likert scale. Results show that 17 out of 24 participants (70.83%) preferred the improved system (exact binomial test, p = 0.043, h = 0.43), and all three perceptual constructs were rated significantly higher for the improved configuration after Holm correction, with large to very large effect sizes (p < 0.001). These findings confirm that the identified technical improvements are perceptible to users in direct interaction and underscore the importance of complementing benchmark evaluation with user-centred evidence when assessing robotic manipulation pipelines.

人机交互机器人抓取用户体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。