arXiv:2606.20905cs.ROcs.AI2026-06被引 2

一个通用机器人模型,统一解决定位、导航与长程规划问题。

Vesta: A Generalist Embodied Reasoning Model

论文配图:Vesta: A Generalist Embodied Reasoning Model
图 1 · 摘自论文原文
  • 用大规模语料和简单记忆机制,让模型具备空间理解与长时间推理能力。
  • 在多个基准测试中,性能比单一任务顶尖模型平均高20%以上。
  • 真实机器人任务成功率提升35%以上,适合需要长期记忆的复杂场景。

开放世界环境中运行的机器人需无缝整合定位、空间推理、导航与长程规划能力。尽管专用模型在单个任务上表现优异,但多模型组合计算开销大且易产生级联错误。我们提出Vesta,一个统一的具身通用模型,将这些能力整合到单一基础模型中。方法结合多样且大规模的精选语料以促进空间定位,并采用简单的多模态记忆机制实现长时间推理。在多个基准测试中,Vesta平均优于单一任务最先进基线超过20%,且超越各领域最优基线的集成方案超过10%——证明通用模型可媲美甚至超越专用模型。在需记忆与推理的真实机器人任务中,Vesta使任务成功率提升超过35%。本工作表明,单一通用模型是比组合专用模型更可行、可扩展且更优的替代方案。

原文摘要 · Abstract (English)

Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at individual tasks, deploying a multi-model stack is computationally expensive and prone to cascading errors. We present Vesta, a unified embodied generalist that consolidates these capabilities into a single foundation model. Our approach combines a diverse and massive curated corpus designed to induce spatial grounding and a simple multimodal memory harness that enables reasoning over extended time horizons. Across diverse benchmarks, Vesta on average beats individual SOTA baselines by >$20\%$ and beats an ensemble of per-category-best baselines by $>10\%$ -- thus demonstrating that a generalist model can match or exceed specialists. On real-world robotic tasks requiring memory and reasoning, Vesta improves task success by >35\%. Our work thus demonstrates that a single generalist is a feasible, scalable, and arguably preferable alternative to combining specialists.

具身智能通用模型机器人长程规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。