arXiv:2608.10756cs.ROcs.CV2026-08中稿 · ACM Multimedia 202…

用3D语义高斯溅射实现机器人在复杂环境中的精准语言引导操作。

Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting

论文配图:Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting
图 1 · 摘自论文原文
  • 通过主动多视角3D语义高斯溅射构建可更新的环境语义地图
  • 真实机器人测试中长程成功率60%,远超现有方法
  • 适合复杂、遮挡多的家用场景下的开放词汇操作任务

具身移动操作需对齐语言、视觉观测、三维场景结构及动作可行性。本文研究局部家庭场景中少样本开放词汇目标定位与操作,提出一个融合主动多视角语义3D高斯溅射(Semantic-3DGS)、可达性感知基座定位和基于扩散的视觉-语言-动作策略的具身多模态定位框架。任务驱动的局部语义3DGS作为跨主动感知、语言条件3D定位、障碍物感知推理、基座准备和动作模型语义调节的共享接口。为保留预训练动作先验,3D语义线索仅注入晚期动作专家模块。在50次真实机器人评估中,全系统长时程成功率达60%,优于PointVLA的40%和DexVLA的28%;在高度杂乱场景中达74%成功率,显著高于单视图变体的52%和PointVLA的46%。系统在75 cm高度偏移下仍保持75%成功率,且消除图像诱导的误抓取。结果表明,显式、可刷新的3D语义定位能有效提升在杂乱、遮挡、视角变化及具身约束下的鲁棒性。

原文摘要 · Abstract (English)

Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before execution. We study open-vocabulary target grounding with few-shot manipulation in local household workspaces and present an embodied multimodal grounding framework that integrates active multi-view Semantic 3D Gaussian Splatting (Semantic-3DGS), reachability-aware base positioning, and a diffusion-based vision-language-action policy. A task-driven local Semantic-3DGS serves as a shared interface across active sensing, language-conditioned 3D localization, obstacle-aware scene reasoning, base preparation, and semantic conditioning of the action model. To preserve pretrained action priors, the 3D semantic cues are injected only into the late action-expert blocks. In expanded 50-trial real-robot evaluations against representative vision-language-action (VLA) approaches, the full system achieves 60% long-horizon success compared with 40% for PointVLA and 28% for DexVLA, and reaches 74% success in heavily cluttered manipulation compared with 52% for the single-view variant and 46% for PointVLA. It also maintains 75% success under a 75 cm height shift and eliminates photo-induced false grasps. These results indicate that explicit, refreshable 3D semantic grounding can improve robustness under clutter, occlusion, viewpoint variation, and embodiment constraints.

具身智能3D语义多模态机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。