arXiv:2503.19281cs.ROcs.AI2025-03被引 7

用视觉语言模型让机器人自主解魔方,突破传统算法局限。

CubeRobot: Grounding Language in Rubik's Cube Manipulation via Vision-Language Model

  • 基于双循环视觉思维链与记忆流,实现多层级任务规划
  • 高阶任务准确率达80%,中低阶任务达100%
  • 适用于需要空间推理与多模态理解的智能体任务

证明魔方定理代表人类高级空间想象与逻辑推理的重要里程碑。传统魔方机器人依赖复杂视觉系统和固定算法,难以适应复杂动态场景。为此,我们提出CubeRobot,一种专用于求解3×3魔方的新型视觉语言模型(VLM),赋予具身智能体多模态理解与执行能力。采用包含43个子任务的CubeCoT图像数据集,涵盖人类难以处理的多种魔方状态。引入双循环视觉思维链架构与记忆流机制,从VLM生成的规划查询中提取任务相关特征,实现独立规划、决策、反思及高低层任务的分离管理。在低层魔方复原任务中,准确率高达100%;中层任务同样达到100%;高层任务准确率为80%。

原文摘要 · Abstract (English)

Proving Rubik's Cube theorems at the high level represents a notable milestone in human-level spatial imagination and logic thinking and reasoning. Traditional Rubik's Cube robots, relying on complex vision systems and fixed algorithms, often struggle to adapt to complex and dynamic scenarios. To overcome this limitation, we introduce CubeRobot, a novel vision-language model (VLM) tailored for solving 3x3 Rubik's Cubes, empowering embodied agents with multimodal understanding and execution capabilities. We used the CubeCoT image dataset, which contains multiple-level tasks (43 subtasks in total) that humans are unable to handle, encompassing various cube states. We incorporate a dual-loop VisionCoT architecture and Memory Stream, a paradigm for extracting task-related features from VLM-generated planning queries, thus enabling CubeRobot to independent planning, decision-making, reflection and separate management of high- and low-level Rubik's Cube tasks. Furthermore, in low-level Rubik's Cube restoration tasks, CubeRobot achieved a high accuracy rate of 100%, similar to 100% in medium-level tasks, and achieved an accuracy rate of 80% in high-level tasks.

魔方机器人视觉语言模型空间推理具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。