arXiv:2604.04811cs.ROcs.CV2026-04中稿 · IEEE Transactions …

用户只需画图加语言,就能指挥家用机器人完成任务。

AnyUser: Translating Sketched User Intent into Domestic Robots

  • 用草图+语言融合理解指令,生成可执行动作
  • 真实机器人实测成功率85.7%~96.4%,跨场景稳定
  • 适合老人、非专业用户,交互效率显著提升

我们提出AnyUser,一种通过在摄像头图像上自由绘制草图(可选语言)实现直观家用任务指令的统一机器人系统。该系统将多模态输入(草图、视觉、语言)解析为空间-语义基本单元,生成无需预先地图或模型的可执行机器人动作。创新点包括多模态融合理解机制和分层策略以增强动作生成鲁棒性。通过三项验证:(1) 在大规模数据集上的定量基准测试显示,系统对多样草图指令具有高解析准确率;(2) 在两台真实机器人平台——固定式7-DoF助手机器臂(KUKA LBR iiwa)与双臂移动操作机(Realman RMC-AIDAL)上完成靶向擦拭、区域清洁等典型任务,证明其在物理环境中可靠落地能力;(3) 涵盖老年群体、模拟非言语者及低技术素养人群的综合用户研究显示,任务完成率高达85.7%-96.4%,用户满意度高。AnyUser弥合了先进机器人能力与非专家用户交互需求之间的鸿沟,为适配真实人类环境的实用助手机器人奠定基础。

原文摘要 · Abstract (English)

We introduce AnyUser, a unified robotic instruction system for intuitive domestic task instruction via free-form sketches on camera images, optionally with language. AnyUser interprets multimodal inputs (sketch, vision, language) as spatial-semantic primitives to generate executable robot actions requiring no prior maps or models. Novel components include multimodal fusion for understanding and a hierarchical policy for robust action generation. Efficacy is shown via extensive evaluations: (1) Quantitative benchmarks on the large-scale dataset showing high accuracy in interpreting diverse sketch-based commands across various simulated domestic scenes. (2) Real-world validation on two distinct robotic platforms, a statically mounted 7-DoF assistive arm (KUKA LBR iiwa) and a dual-arm mobile manipulator (Realman RMC-AIDAL), performing representative tasks like targeted wiping and area cleaning, confirming the system's ability to ground instructions and execute them reliably in physical environments. (3) A comprehensive user study involving diverse demographics (elderly, simulated non-verbal, low technical literacy) demonstrating significant improvements in usability and task specification efficiency, achieving high task completion rates (85.7%-96.4%) and user satisfaction. AnyUser bridges the gap between advanced robotic capabilities and the need for accessible non-expert interaction, laying the foundation for practical assistive robots adaptable to real-world human environments.

人机交互草图理解家用机器人多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。