构建交互式推理新基准,评估机器人在真实场景中的长期任务规划能力。
Chain Of Interaction Benchmark (COIN): When Reasoning meets Embodied Interaction

- 设计50个日常任务,分基础、组合两类,模拟连续交互与因果推理。
- 收集1000条示范数据,通过低成本移动AR系统实现高精度动作捕捉。
- 揭示现有模型在视觉理解与执行间的巨大差距,适合研究具身智能者参考。
通用具身智能体需在部分可观测环境下持续与环境交互,进行因果依赖性推理,以完成长时程任务,才能应用于真实场景。例如从柜子中取苹果,需依次打开多扇门与抽屉,直至目标可见且可及,要求顺序交互与动态规划。然而,现有基准未能系统评估此核心能力。本文提出COIN基准,包含三个关键贡献:其一,构建COIN-50(50个日常交互任务),并细分出需因果依赖的COIN-Primitive和中等复杂度的COIN-Composition,用于技能学习与泛化评估;其二,开发低成本移动AR远程操控系统,采集每项基础任务50条示范数据(共1000条);其三,建立执行稳定性与泛化鲁棒性的系统评估指标,对CodeAsPolicy、VLA及语言条件下的H-VLA方法进行评测。全面评估发现当前模型在交互推理任务上表现不佳,主要源于视觉理解与运动执行之间的显著鸿沟,并提供细粒度分析。
原文摘要 · Abstract (English)
Generalist embodied agents must perform interactive, causally-dependent reasoning, continually interacting with the environment, acquiring information, and updating plans to solve long-horizon tasks before they could be adopted in real-life scenarios. For instance, retrieving an apple from a cabinet may require opening multiple doors and drawers before the apple becomes visible and reachable, demanding sequential interaction under partial observability. However, existing benchmarks fail to systematically evaluate this essential capability. We introduce COIN, a benchmark designed to assess interactive reasoning in realistic robotic manipulation through three key contributions. First, we construct COIN-50: 50 interactive tasks in daily scenarios, and create COIN-Primitive required by causally-dependent tasks, and COIN-Composition with mid-term complexity for skill learning and generalization evaluation. Second, we develop a low-cost mobile AR teleoperation system and collect the COIN-Primitive Dataset with 50 demonstrations per primitive task (1,000 in total). Third, we develop systematic evaluation metrics about execution stability and generalization robustness to evaluate CodeAsPolicy, VLA, and language-conditioned H-VLA approaches. Our comprehensive evaluation reveals critical limitations in current methods: models struggle with interactive reasoning tasks due to significant gaps between visual understanding and motor execution. We provide fine-grained analysis of these limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。