让AI在真实环境中像人一样边看边想边行动,解决复杂交互任务。
Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
- 用视觉-思考-动作三阶段流程模拟人类互动决策过程。
- 在长周期任务中比o1、Claude-3.7等模型高13%~24%成功率。
- 适合需要持续观察与反思的机器人导航、智能助手等场景。
深度思维模型在数学与编程任务中表现优异,但在需通过图像与动作交替进行连续交互的具身领域仍缺乏探索。本文提出Embodied Reasoner,将o1类推理范式扩展至具身交互搜索任务。不同于依赖逻辑推导的数学推理,具身场景要求空间理解、时间推理及基于交互历史的持续自省。为此,我们构建了包含9.3k个连贯观察-思考-动作轨迹的数据集,含64k张交互图像与90k种多样化的思维过程(分析、空间推理、反思、规划、验证)。设计三阶段训练流程:模仿学习、通过拒绝采样实现自我探索、通过反思调优进行自我修正。评估显示,该模型显著优于现有先进视觉推理模型,相较OpenAI o1、o3-mini、Claude-3.7分别提升+9%、+24%、+13%。分析表明,模型重复搜索与逻辑矛盾更少,尤其在复杂长时任务中优势明显;真实环境测试亦验证其优越性,且重复搜索与逻辑不一致案例更少。
原文摘要 · Abstract (English)
Recent advances in deep thinking models have demonstrated remarkable reasoning capabilities on mathematical and coding tasks. However, their effectiveness in embodied domains which require continuous interaction with environments through image action interleaved trajectories remains largely -unexplored. We present Embodied Reasoner, a model that extends o1 style reasoning to interactive embodied search tasks. Unlike mathematical reasoning that relies primarily on logical deduction, embodied scenarios demand spatial understanding, temporal reasoning, and ongoing self-reflection based on interaction history. To address these challenges, we synthesize 9.3k coherent Observation-Thought-Action trajectories containing 64k interactive images and 90k diverse thinking processes (analysis, spatial reasoning, reflection, planning, and verification). We develop a three-stage training pipeline that progressively enhances the model's capabilities through imitation learning, self-exploration via rejection sampling, and self-correction through reflection tuning. The evaluation shows that our model significantly outperforms those advanced visual reasoning models, e.g., it exceeds OpenAI o1, o3-mini, and Claude-3.7 by +9\%, 24\%, and +13\%. Analysis reveals our model exhibits fewer repeated searches and logical inconsistencies, with particular advantages in complex long-horizon tasks. Real-world environments also show our superiority while exhibiting fewer repeated searches and logical inconsistency cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。