arXiv:2606.17480cs.CVcs.RO2026-06

提升机器人规划的3D重建与记忆管理,让动作更准更可靠

GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning

论文配图:GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning
图 1 · 摘自论文原文
  • 用多视角几何先验修正单目3D重建,避免幻觉
  • 重建精度提升:CD降低2.20%,PSNR提高2.36%
  • 记忆系统可控制质量与冲突,适合复杂任务规划

通用视觉-语言-动作系统需要以对象为中心的3D证据和可复用的操作经验来规划可靠的机器人轨迹。GeneralVLA 提供了将语言和RGB-D观测转化为3D末端执行器路径的分层接口,但仍存在两大瓶颈。首先,单目SAM3D式物体重建可能产生姿态幻觉和未见几何,而多视角观测在标定后能提供更稳定的物体形状。其次,原始KnowledgeBank主要检索语义相似片段并追加新知识,难以控制记忆质量、冲突、置信度和几何相关性。为解决第一挑战,我们提出GeoFuse-MV3D,一个基于几何先验的多视角SAM3D重建分支,通过输入视图掩码验证外部几何线索,应用软视觉壳支持,进行轴向精修,并仅融合几何信息而保留外观特征。为解决第二挑战,我们将KnowledgeBank升级为带显式质量、置信度、生命周期、验证器和冲突元数据的受控长期记忆系统,配合精准检索。最终,在GSO-30上评估重建分支,在Terminal-Bench 2.0和SWE-Bench Verified上评估记忆模块;GeoFuse-MV3D相比MV-SAM3D基线,使CD和LPIPS分别降低2.20%和2.02%,同时提升PSNR和SSIM 2.36%和1.03%;KnowledgeBank相比ReasoningBank,在Terminal-Bench SR提升4.53%,在SWE-Bench解析率提升3.73%,同时减少AS 4.95%和5.65%。

原文摘要 · Abstract (English)

Generalist vision-language-action systems need object-centric 3D evidence and reusable manipulation experience to plan reliable robot trajectories. GeneralVLA provides a hierarchical interface for converting language and RGB-D observations into 3D end-effector paths, but two bottlenecks remain. First, monocular SAM3D-style object reconstruction can hallucinate pose and unseen geometry, while manipulation benefits from stable object shape when calibrated multi-view observations are available. Second, the original KnowledgeBank mainly retrieves semantically similar snippets and appends new knowledge, which makes it difficult to control memory quality, conflicts, confidence, and geometric relevance. To address the first challenge, we introduce GeoFuse-MV3D, a geometry-prior-guided MV-SAM3D reconstruction branch that verifies external geometry cues with input-view masks, applies soft visual-hull support, performs axis-wise refinement, and fuses only geometry while preserving appearance. To address the second challenge, we upgrade KnowledgeBank into a governed long-term memory system with explicit quality, confidence, lifecycle, verifier, and conflict metadata, together with precision-oriented retrieval. Finally, we evaluate the reconstruction branch on GSO-30 and the memory module on Terminal-Bench 2.0 and SWE-Bench Verified; GeoFuse-MV3D improves over the MV-SAM3D baseline by reducing CD and LPIPS by 2.20% and 2.02% while increasing PSNR and SSIM by 2.36% and 1.03%, and KnowledgeBank improves over ReasoningBank by 4.53% on Terminal-Bench SR and 3.73% on SWE-Bench resolve rate, while reducing AS by 4.95% and 5.65%, respectively. Code: https://github.com/AIGeeksGroup/GeneralVLA-2. Website: https://aigeeksgroup.github.io/GeneralVLA-2.

机器人规划3D重建记忆系统多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。