构建机器人操作与具身世界建模的统一中间表征体系
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

- 提供超过23万次操作片段的密集标注数据集
- 支持空间时间联合推理与长时序物理状态预测
- 适合研究具身智能、机器人规划与仿真建模的学者
现有机器人数据集仍存在成本高、具身依赖性强、细粒度结构标注不足等问题,难以支撑通用推理、执行及长时序环境动态模拟。基于前期工作RoboInter1.0,本文提出RoboInter1.5,一个涵盖机器人操作与具身世界建模的综合性中间表征套件。RoboInter-Data包含超过230,000次操作片段,覆盖571个场景,每帧均有超过十类中间表示的密集标注,包括子任务、基础技能、物体与夹爪定位、分割、可操作性、抓取姿态、接触点、运动轨迹等。基于此,RoboInter-VQA引入时空具身视觉问答任务,用于评估与提升RoboInter-VLM的中间表征推理能力;RoboInter-VLA进一步研究此类表征在隐式、显式与模块化计划-执行范式下的动作执行效益。为更好建模物理世界,我们引入RoboInter-World,利用中间表征作为结构化条件信号,实现对未来世界状态的可控预测。大量实验表明,RoboInter1.5为中间表征提供了统一的时空框架。其不仅将中间表征视为可解释信号,更将其构建成双向接口,既能规整低层动作空间,又能约束开放世界物理模拟器的潜在演化路径。
原文摘要 · Abstract (English)
Existing robot datasets remain expensive to curate, embodiment-specific, and insufficiently annotated with the fine-grained structure required for generalizable reasoning, execution, or long-horizon environment dynamics simulation. Building on our prior work, RoboInter1.0, we present RoboInter1.5, an extended and holistic suite of intermediate representations for both robotic manipulation and embodied world modeling. RoboInter1.5 provides a unified resource of data, benchmarks, and models centered on dense manipulation-oriented intermediate representations. Specifically, RoboInter-Data contains over 230k manipulation episodes across 571 scenes with dense per-frame annotations covering more than ten types of intermediate representations, including subtasks, primitive skills, object and gripper grounding, segmentation, affordance, grasp poses, contact points, motion traces, etc. Built upon these annotations, RoboInter-VQA introduces spatial and temporal embodied VQA tasks to benchmark and improve the intermediate-representation reasoning capabilities of our RoboInter-VLM. RoboInter-VLA further studies how such representations benefit action execution through implicit, explicit, and modular plan-then-execute paradigms. To better model the physical world, we further introduce RoboInter-World, which leverages intermediate representations as structured conditioning signals for controllable prediction of future world states. Extensive evaluations demonstrate that RoboInter1.5 provides a unified spatiotemporal scaffolding for intermediate representations. Rather than treating intermediate representations merely as interpretable signals, RoboInter1.5 conceptualizes them as a bidirectional interface that both regularizes low-level action spaces and constrains the latent rollouts of open-world physical simulators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。