arXiv:2602.09973cs.RO2026-02被引 15

构建机器人操作中间表示体系,提升视觉语言模型泛化能力。

RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation

  • 提出统一中间表示框架,支持从高阶规划到低阶执行的全流程建模。
  • 构建超23万条轨迹数据集,覆盖571场景与10类以上中间表示,规模领先。
  • 适配多模态大模型研究者,尤其适合关注具身推理与可泛化机器人学习人群。

大型视觉语言模型(VLM)的发展推动了视觉-语言-动作(VLA)系统在机器人操作中的应用。然而,现有操作数据集仍存在标注成本高、高度依赖具体机器人本体、覆盖范围与多样性不足等问题,制约了VLA模型的泛化能力。近期方法尝试通过“先规划再执行”范式缓解上述问题,即先生成高层计划(如子任务、动作序列),再转化为底层动作,但这类方法严重依赖额外的中间监督信号,而现有数据集普遍缺乏此类信息。为此,我们提出RoboInter操作套件,一个包含数据、基准和模型的统一资源平台,涵盖中间表示的完整生态。其中,RoboInter-Tool是一个轻量级图形界面工具,支持多样化表示的半自动标注;RoboInter-Data是一个大规模数据集,包含超过23万条轨迹,覆盖571个多样场景,提供超过10类中间表示的密集帧级标注,显著超越以往工作在规模与标注质量上的表现。基于此,RoboInter-VQA引入9类空间与20类时间维度的具身视觉问答任务,系统性地评估并提升VLM的具身推理能力。同时,RoboInter-VLA提供一体化的“先规划后执行”框架,支持模块化与端到端的VLA变体,通过中间监督实现高层规划与底层执行的高效衔接。总体而言,RoboInter为通过细粒度、多样化的中间表示推进鲁棒且可泛化的机器人学习提供了实用基础。

原文摘要 · Abstract (English)

Advances in large vision-language models (VLMs) have stimulated growing interest in vision-language-action (VLA) systems for robot manipulation. However, existing manipulation datasets remain costly to curate, highly embodiment-specific, and insufficient in coverage and diversity, thereby hindering the generalization of VLA models. Recent approaches attempt to mitigate these limitations via a plan-then-execute paradigm, where high-level plans (e.g., subtasks, trace) are first generated and subsequently translated into low-level actions, but they critically rely on extra intermediate supervision, which is largely absent from existing datasets. To bridge this gap, we introduce the RoboInter Manipulation Suite, a unified resource including data, benchmarks, and models of intermediate representations for manipulation. It comprises RoboInter-Tool, a lightweight GUI that enables semi-automatic annotation of diverse representations, and RoboInter-Data, a large-scale dataset containing over 230k episodes across 571 diverse scenes, which provides dense per-frame annotations over more than 10 categories of intermediate representations, substantially exceeding prior work in scale and annotation quality. Building upon this foundation, RoboInter-VQA introduces 9 spatial and 20 temporal embodied VQA categories to systematically benchmark and enhance the embodied reasoning capabilities of VLMs. Meanwhile, RoboInter-VLA offers an integrated plan-then-execute framework, supporting modular and end-to-end VLA variants that bridge high-level planning with low-level execution via intermediate supervision. In total, RoboInter establishes a practical foundation for advancing robust and generalizable robotic learning via fine-grained and diverse intermediate representations.

机器人操作中间表示具身智能视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。