让机器人能听懂模糊指令并动态调整行为,实现自然协作。
Ludi${}_{\scriptscriptstyle 0.1}$: An Agentic System for Socially Intelligent Robots

- 用多轮对话数据训练视觉语言模型,理解复杂社交与任务混合请求。
- 集成语音、记忆、导航与操作策略,支持上下文连续互动。
- 适合研究人机协作、具身智能的开发者与研究人员参考。
机器人基础模型在感知与控制方面已取得显著进展,但自然的人机协作不仅需要执行孤立指令。机器人还需识别意图模糊性,保持跨轮次对话上下文,清晰表达自身意图,并根据用户意图变化动态调整行为。我们提出 $ exttt{Ludi}_{ exttt{0.1}}$,一个面向社交智能机器人的智能体系统,整合交互式语音、多模态推理、记忆、导航与学习型操作能力。其决策核心为基于多轮交互轨迹微调的视觉语言模型,涵盖模糊请求、澄清、修正、打断、社交与任务混合对话及多步任务。专用框架管理模型与工具的交互循环,特定导航与操作策略执行物理技能。Ludi${}_{ exttt{0.1}}$ 为当前实现流畅人机协作提供了可行路径,同时生成了构建更深度集成的机器人-人类基础模型所需的多模态交互数据。
原文摘要 · Abstract (English)
Robot foundation models have substantially advanced perception and control, but natural human-robot collaboration requires more than executing isolated commands. A robot must recognize ambiguity, maintain context across turns, communicate its intentions, and revise ongoing behavior as the user's intent changes. We present $\scriptstyle\mathsf{Ludi}_{\scriptscriptstyle 0.1}$, an agentic system for socially intelligent robots that integrates interactive speech, multimodal reasoning, memory, navigation, and learned manipulation. Its decision-making core is a fine-tuned vision-language model trained on multi-turn interaction traces spanning ambiguous requests, clarifications, corrections, interruptions, mixed social and task dialogue, and multi-step tasks. A purpose-built harness manages the model-tool interaction loop, while specialized navigation and manipulation policies execute physical skills. Ludi${}_{\scriptscriptstyle 0.1}$ demonstrates a practical path toward fluid human-robot collaboration today while producing the multimodal interaction traces needed to develop a more deeply integrated foundation model for robots and people.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。