arXiv:2510.24161cs.AIcs.MM2025-10被引 1

一个能跨空间、跨任务、跨机器人形态通用的大型模型,让AI在虚实世界中都能智能行动。

BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning

  • 通过两阶段训练,将实体知识注入大模型,保持语言理解能力
  • 单个模型在数字与物理任务上分别提升约6%和3%表现
  • 适合需要跨平台、多机器人协同的智能系统研发者

多模态大语言模型(MLLM)虽推动了视觉-语言推理发展,并广泛部署于具身智能体中,但仍存在显著局限:跨数字-物理空间泛化能力差;视觉-语言-动作模型(VLAs)仅生成底层动作,缺乏稳健的高层具身推理;多数具身大语言模型(ELLM)局限于数字空间,难以迁移到真实世界。因此,能够在数字与物理空间间无缝运行,并实现跨形态、跨任务泛化的统一模型仍属空白。本文提出‘无界大模型(BLM₁)’,一种多模态空间基础模型,兼具指令遵循、推理能力与具身知识,支持鲁棒的跨形态控制。该模型通过两阶段训练实现三大能力:跨空间迁移、跨任务学习与跨形态泛化。第一阶段利用精心构建的数字语料库向MLLM注入具身知识,同时保持语言能力;第二阶段通过意图桥接接口,从MLLM提取高层语义并指导控制策略,无需微调模型主干。该过程基于自收集的跨形态示范数据集,涵盖四类机器人形态与六项逐步挑战的任务。在数字与物理基准上的评估显示,单一BLM₁实例在数字任务上优于四个模型家族(MLLMs、ELLMs、VLAs、GMLMs)约6%,在物理任务上提升约3%。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs generalize poorly across digital-physical spaces and embodiments; vision-language-action models (VLAs) produce low-level actions yet lack robust high-level embodied reasoning; and most embodied large language models (ELLMs) are constrained to digital-space with poor generalization to the physical world. Thus, unified models that operate seamlessly across digital and physical spaces while generalizing across embodiments and tasks remain absent. We introduce the \textbf{Boundless Large Model (BLM$_1$)}, a multimodal spatial foundation model that preserves instruction following and reasoning, incorporates embodied knowledge, and supports robust cross-embodiment control. BLM$_1$ integrates three key capabilities -- \textit{cross-space transfer, cross-task learning, and cross-embodiment generalization} -- via a two-stage training paradigm. Stage I injects embodied knowledge into the MLLM through curated digital corpora while maintaining language competence. Stage II trains a policy module through an intent-bridging interface that extracts high-level semantics from the MLLM to guide control, without fine-tuning the MLLM backbone. This process is supported by a self-collected cross-embodiment demonstration suite spanning four robot embodiments and six progressively challenging tasks. Evaluations across digital and physical benchmarks show that a single BLM$_1$ instance outperforms four model families -- MLLMs, ELLMs, VLAs, and GMLMs -- achieving $\sim\!\textbf{6%}$ gains in digital tasks and $\sim\!\textbf{3%}$ in physical tasks.

具身智能多模态模型跨空间机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。