用打结的绳子测试智能体的空间推理与操作能力。
Knot So Simple: A Minimalistic Environment for Spatial Reasoning
- 基于绳结交叉数设计渐进复杂度任务,纯视觉输入下挑战空间推理。
- 多类方法在复杂任务中表现受限,暴露感知-推理-操作融合难题。
- 适合研究视觉导航、具身智能和复杂决策的科研人员使用。
我们提出KnotGym,一个用于复杂空间推理与操作的交互式环境。KnotGym包含一系列目标导向的绳索操作任务,复杂度随绳结交叉数递增,所有任务均需仅从图像观测中进行决策。该环境提供清晰可量化的复杂度轴,构成自然的泛化测试基准。其简洁的观测空间支持可扩展开发,同时凸显了精准感知、空间推理与具身操作融合的核心挑战。我们评估了模型基强化学习、模型预测控制及链式思维推理等不同类别的方法,揭示了KnotGym带来的关键难题。KnotGym已开源,地址为https://github.com/lil-lab/knotgym。
原文摘要 · Abstract (English)
We propose KnotGym, an interactive environment for complex, spatial reasoning and manipulation. KnotGym includes goal-oriented rope manipulation tasks with varying levels of complexity, all requiring acting from pure image observations. Tasks are defined along a clear and quantifiable axis of complexity based on the number of knot crossings, creating a natural generalization test. KnotGym has a simple observation space, allowing for scalable development, yet it highlights core challenges in integrating acute perception, spatial reasoning, and grounded manipulation. We evaluate methods of different classes, including model-based RL, model-predictive control, and chain-of-thought reasoning, and illustrate the challenges KnotGym presents. KnotGym is available at https://github.com/lil-lab/knotgym.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。