arXiv:2604.04664cs.ROcs.AI2026-04

让不同机器人协作更智能,用统一模型打通语言理解与物理执行

ROSClaw: A Hierarchical Semantic-Physical Framework for Heterogeneous Multi-Agent Collaboration

  • 用统一视觉语言模型控制异构机器人,实现语义与动作无缝衔接
  • 支持真实世界数据回传与迭代优化,提升多智能体任务鲁棒性
  • 适合需要跨平台协作的机器人研发团队快速部署和持续改进

将大语言模型与具身智能体结合虽提升了高层推理能力,但语义理解与物理执行之间仍存在关键鸿沟。现有视觉-语言-动作(VLA)和视觉-语言导航(VLN)系统在长时序、结构化任务中表现不足,且依赖模块化流水线,导致实验验证与策略优化成本高昂。为此,我们提出ROSClaw框架,通过统一的视觉语言模型(VLM)控制器,集成策略学习与任务执行,利用e-URDF表示异构机器人的物理约束,构建仿真到现实的拓扑映射,实现实时访问模拟与真实智能体的物理状态。框架还引入数据收集与状态累积机制,在真实执行中存储机器人状态、多模态观测与执行轨迹,支持后续迭代优化。部署时,统一智能体保持语义连续性,并动态分配任务控制给不同智能体,增强多策略执行鲁棒性。通过建立自主闭环系统,ROSClaw降低对机器人定制开发流程的依赖,支持硬件级验证、自动化的SDK级控制程序生成与工具化执行,实现技能的快速跨平台迁移与持续进化。

原文摘要 · Abstract (English)

The integration of large language models (LLMs) with embodied agents has improved high-level reasoning capabilities; however, a critical gap remains between semantic understanding and physical execution. While vision-language-action (VLA) and vision-language-navigation (VLN) systems enable robots to perform manipulation and navigation tasks from natural language instructions, they still struggle with long-horizon sequential and temporally structured tasks. Existing frameworks typically adopt modular pipelines for data collection, skill training, and policy deployment, resulting in high costs in experimental validation and policy optimization. To address these limitations, we propose ROSClaw, an agent framework for heterogeneous robots that integrates policy learning and task execution within a unified vision-language model (VLM) controller. The framework leverages e-URDF representations of heterogeneous robots as physical constraints to construct a sim-to-real topological mapping, enabling real-time access to the physical states of both simulated and real-world agents. We further incorporate a data collection and state accumulation mechanism that stores robot states, multimodal observations, and execution trajectories during real-world execution, enabling subsequent iterative policy optimization. During deployment, a unified agent maintains semantic continuity between reasoning and execution, and dynamically assigns task-specific control to different agents, thereby improving robustness in multi-policy execution. By establishing an autonomous closed-loop framework, ROSClaw minimizes the reliance on robot-specific development workflows. The framework supports hardware-level validation, automated generation of SDK-level control programs, and tool-based execution, enabling rapid cross-platform transfer and continual improvement of robotic skills. Ours project page: https://www.rosclaw.io/.

多智能体机器人控制大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。