用新框架提升机器人在复杂任务中的推理与精准动作执行能力
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
- 提出分步评估机器人推理能力的ERIQ基准,覆盖6000+问题
- 设计离散动作编码器FACT,实现高保真轨迹重建并提升真实场景表现
- 统一推理与动作空间,兼顾泛化性与控制精度,适合通用机器人研发
开放世界中通用机器人系统需同时具备广泛泛化能力和高精度动作执行,而现有视觉-语言-动作(VLA)模型难以兼顾。尽管大型视觉-语言模型(VLM)提升了语义泛化能力,但缺乏具身推理导致行为脆弱;仅强推理又无法实现精确控制。为此,我们提出具身推理智商(ERIQ),一个大规模具身推理基准,涵盖4个推理维度、6000+问答对。通过解耦推理与执行,ERIQ揭示了具身推理能力与端到端VLA泛化之间存在强正相关。为弥合推理到精确执行的差距,我们提出基于流匹配的动作分词器FACT,将连续控制转化为离散序列,同时保持高保真轨迹重建。由此构建的GenieReasoner在统一空间中联合优化推理与动作,在真实任务中超越连续与先前离散动作基线。ERIQ与FACT共同构成诊断和突破推理-精度权衡的系统性框架,推动鲁棒、通用的机器人操作发展。
原文摘要 · Abstract (English)
General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challenging for existing Vision-Language-Action (VLA) models. While large Vision-Language Models (VLMs) improve semantic generalization, insufficient embodied reasoning leads to brittle behavior, and conversely, strong reasoning alone is inadequate without precise control. To provide a decoupled and quantitative assessment of this bottleneck, we introduce Embodied Reasoning Intelligence Quotient (ERIQ), a large-scale embodied reasoning benchmark in robotic manipulation, comprising 6K+ question-answer pairs across four reasoning dimensions. By decoupling reasoning from execution, ERIQ enables systematic evaluation and reveals a strong positive correlation between embodied reasoning capability and end-to-end VLA generalization. To bridge the gap from reasoning to precise execution, we propose FACT, a flow-matching-based action tokenizer that converts continuous control into discrete sequences while preserving high-fidelity trajectory reconstruction. The resulting GenieReasoner jointly optimizes reasoning and action in a unified space, outperforming both continuous-action and prior discrete-action baselines in real-world tasks. Together, ERIQ and FACT provide a principled framework for diagnosing and overcoming the reasoning-precision trade-off, advancing robust, general-purpose robotic manipulation. Project page: https://geniereasoner.github.io/GenieReasoner/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。