arXiv:2601.16870cs.RO2026-01中稿 · IEEE RAS/EMBS 11th…

构建对话式助老机器人多模态数据集,解决交互模糊问题。

A Multimodal Data Collection Framework for Dialogue-Driven Assistive Robotics to Clarify Ambiguities: A Wizard-of-Oz Pilot Study

  • 用双室巫师实验模拟机器人自主,收集自然对话交互
  • 采集53组数据,涵盖5类任务与5种同步模态信号
  • 适合研究对话驱动的智能辅助控制与模糊处理

轮椅与轮椅安装机械臂(WMRA)的集成控制对严重运动障碍者具有提升独立性的潜力,但现有界面常缺乏直观交互所需的灵活性。尽管数据驱动的AI方法前景广阔,但进展受限于缺乏捕捉自然人机交互(HRI)的多模态数据集,尤其在对话式控制中的语义模糊性。为此,我们提出一种多模态数据采集框架,采用基于对话的交互协议和两室巫师实验(WoZ)设置,在模拟机器人自主的同时诱发自然用户行为。该框架同步记录五种模态:RGB-D视频、对话音频、惯性测量单元(IMU)信号、末端执行器笛卡尔位姿及全身关节状态,覆盖五类助老任务。基于此框架,我们从五名参与者处收集了53次试验数据,并通过运动平滑性分析与用户反馈验证了数据质量。结果表明,该框架能有效捕捉多样化的模糊类型,支持自然的对话式交互,具备扩展为更大规模数据集的潜力,可用于学习、基准测试与评估具模糊感知能力的辅助控制。

原文摘要 · Abstract (English)

Integrated control of wheelchairs and wheelchair-mounted robotic arms (WMRAs) has strong potential to increase independence for users with severe motor limitations, yet existing interfaces often lack the flexibility needed for intuitive assistive interaction. Although data-driven AI methods show promise, progress is limited by the lack of multimodal datasets that capture natural Human-Robot Interaction (HRI), particularly conversational ambiguity in dialogue-driven control. To address this gap, we propose a multimodal data collection framework that employs a dialogue-based interaction protocol and a two-room Wizard-of-Oz (WoZ) setup to simulate robot autonomy while eliciting natural user behavior. The framework records five synchronized modalities: RGB-D video, conversational audio, inertial measurement unit (IMU) signals, end-effector Cartesian pose, and whole-body joint states across five assistive tasks. Using this framework, we collected a pilot dataset of 53 trials from five participants and validated its quality through motion smoothness analysis and user feedback. The results show that the framework effectively captures diverse ambiguity types and supports natural dialogue-driven interaction, demonstrating its suitability for scaling to a larger dataset for learning, benchmarking, and evaluation of ambiguity-aware assistive control.

助老机器人多模态数据对话交互巫师实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。