构建可长期共存的多机器人助手,支持用户界面与任务灵活组合。
A Glimpse into Long-term Physical Coexistence with Intelligent Robots

- 以机器人网关为核心设计统一能力接口,分离高层推理与底层执行。
- 在Astribot S1上验证,支持复杂家务任务如背包打包、垃圾袋提起。
- 适合研究人机协作、多机器人系统集成的开发者与研究人员。
长期与智能机器人共存不仅需要强大机器人策略,还需支持多样用户交互界面、长期记忆用户偏好、跨机器人形态协调,并将人类意图安全转化为物理执行。我们提出PHILIA,一种基于机器人网关抽象的多机器人代理。它保留OpenClaw丰富的交互与工具生态,通过统一接口暴露本地运行时、感知、导航、语音播放及机器人策略。该设计将低频高语义的代理推理与高频低层执行解耦,实现用户界面、机器人形态、策略后端的即插即用集成。用户交互因此具备组合性:任一组件(界面、形态、策略、导航或交互算法)改进,无需重构系统即可提升整体体验。我们在Astribot S1机器人上验证架构,并设计机器人网关合约,支持未来异构平台通过共享能力接口实现观测、任务执行、导航、语音播放、状态监控和任务取消。展示了代理记忆与场景理解扎根于机器人动作的代表性用例,涵盖从简单整理到复杂长周期、灵巧服务任务,如打包背包、提起垃圾袋。强调人机交互流程中,结合上下文理解用户意图与偏好,以及执行中的闭环确认或调整,是有效协助的关键。
原文摘要 · Abstract (English)
Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must support diverse user-facing interfaces, maintain long-horizon memory of people and preferences, coordinate across robot embodiments, and translate human intent into safe physical execution. We introduce PHILIA, a multi-robot agent built around a robot gateway abstraction. PHILIA retains the rich interaction and tool ecosystem of OpenClaw while exposing robot-local runtimes, onboard perception, navigation, speaker, and robot policies through a unified capability interface. This design decouples low-frequency, high-semantic agent reasoning from high-frequency, low-level robot execution, enabling plug-and-play integration of user interfaces, robot embodiments, and policy backends. As a result, the user experience becomes compositional: advances in user interfaces, robot embodiments, robot policies, navigation, or interaction algorithms can improve the overall experience without redesigning the system. We validate the architecture on Astribot S1 robots while designing the robot gateway contract to support future heterogeneous robot platforms through a shared capability interface for observation, task execution, navigation, speech playback, status monitoring, and task cancellation. We present representative use cases in which agent memory and scene understanding are grounded in robot actions. These span interactive household scenarios, ranging from simple organization to challenging long-horizon and dexterous service tasks, such as packing a backpack and lifting a garbage bag. We highlight the human-robot interaction flow, where contextual understanding of user intent and preferences, together with human-in-the-loop confirmation or adjustment during execution, is essential for effective assistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。