arXiv:2607.12894cs.CV2026-07被引 4

高效物理世界智能体模型,支持复杂交互与长时推理。

Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

论文配图:Hy-Embodied-VLM-1.0: Efficient Physical-World Agents
图 1 · 摘自论文原文
  • 以动作为中心构建三阶段能力体系,指导数据与训练设计
  • 38项基准测试中19项领先,32B参数模型性能提升8.4%
  • 激活30亿参数达32B级效果,适合低延迟部署场景

构建具备实际操作能力的具身智能体,不仅需要多模态感知与理解,还需具备推理行动、适应变化和与物理世界交互的智能。本文提出Hy-Embodied-VLM-1.0,一种专为物理世界具身智能体设计的高效基础模型。基于动作中心的能力分类体系(包含动作相关状态理解、动作转移推理、序列与自适应推理三个维度),构建系统化数据流程并混合预训练与后训练数据。模型采用Hy3-A3B语言骨干与Hy-ViT2视觉编码器,结合高效的Mixture-of-Experts架构,在保持高推理效率的同时具备强模型容量。在涵盖38个基准的综合性评估中,该模型在19项任务上优于同类规模模型,显著超越Qwen3.6-A3B和Cosmos 3。相比前代版本Hy-Embodied-0.5 MoT-2B,平均性能提升8.4%。尽管仅激活30亿参数,其表现接近前代320亿参数模型。此外,在需多轮交互与长时推理的具身智能任务中也表现出色。

原文摘要 · Abstract (English)

Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situations, and interacting with the physical world. In this report, we introduce Hy-Embodied-VLM-1.0, an efficient and powerful embodied foundation model specifically designed for embodied agents operating in the physical world. To cultivate such capabilities from the pre-training stage onward, we define an action-centric capability taxonomy comprising three progressive dimensions: Action-Relevant State Understanding, Action-Transition Reasoning, and Sequential and Adaptive Reasoning. Guided by this taxonomy, we develop a systematic data pipeline and curate data mixtures spanning both pre-training and post-training. To deliver strong physical-world understanding and interaction capabilities while supporting latency-sensitive deployment, we build our model on the Hy3-A3B language backbone and the Hy-ViT2 vision encoder. Its efficient Mixture-of-Experts architecture combines strong model capacity with high inference efficiency. We evaluate Hy-Embodied-VLM-1.0 on a comprehensive suite of 38 benchmarks covering embodied perception, physical-world understanding, and embodied reasoning. The model achieves the best performance among similarly sized models on 19 of the 38 benchmarks and substantially outperforms strong competitors, including Qwen3.6-A3B and Cosmos 3. Compared with the previous-generation Hy-Embodied-0.5 MoT-2B, Hy-Embodied-VLM-1.0 improves average performance by 8.4%. Despite activating only 3B parameters, it achieves performance close to that of the previous-generation model with 32B activated parameters. Beyond static benchmark evaluation, Hy-Embodied-VLM-1.0 also demonstrates strong performance on embodied agentic tasks requiring multi-turn interaction and long-horizon reasoning.

具身智能多模态高效推理智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。