arXiv:2604.18484cs.CVcs.MM2026-04被引 1

XEmbodied让视觉语言模型具备3D空间和物理交互能力,提升复杂场景理解。

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments

论文配图:XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments
图 1 · 摘自论文原文
  • 通过3D适配器融合几何信息,用上下文标记提取物理信号。
  • 在18个公开基准上显著提升空间推理与分布外泛化能力。
  • 适合需要大规模场景理解的机器人、自动驾驶等应用。

视觉-语言-动作(VLA)模型推动新一代自主系统发展,但其训练需来自复杂环境的大规模高质量标注。现有云端流水线依赖仅在2D图像-文本上预训练的通用视觉语言模型(VLM),缺乏几何推理与领域语义。为此,我们提出XEmbodied,一种云端基础模型,为VLM注入内在的3D几何感知与物理线索(如占据网格、3D边界框)交互能力。不同于将几何作为辅助输入,XEmbodied通过结构化3D适配器整合几何表示,并利用高效图像-具身适配器将物理信号蒸馏为上下文标记。结合渐进式领域课程与强化学习后训练,XEmbodied在保持通用能力的同时,在18个公开基准上展现稳健性能,显著提升空间推理、交通语义、具身可操作性及分布外泛化能力,适用于大规模场景挖掘与具身视觉问答。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex environments. Current cloud pipelines rely on generic vision-language models (VLMs) that lack geometric reasoning and domain semantics due to their 2D image-text pretraining. To address this mismatch, we propose XEmbodied, a cloud-side foundation model that endows VLMs with intrinsic 3D geometric awareness and interaction with physical cues (e.g., occupancy grids, 3D boxes). Instead of treating geometry as auxiliary input, XEmbodied integrates geometric representations via a structured 3D Adapter and distills physical signals into context tokens using an Efficient Image-Embodied Adapter. Through progressive domain curriculum and reinforcement learning post-training, XEmbodied preserves general capabilities while demonstrating robust performance across 18 public benchmarks. It significantly improves spatial reasoning, traffic semantics, embodied affordance, and out-of-distribution generalization for large-scale scenario mining and embodied VQA.

具身智能3D建模视觉语言模型物理推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。