arXiv:2606.10568cs.RO2026-06被引 2

让机器人在执行动作前,用3D空间信息验证动作对错,提升操作可靠性。

VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models

论文配图:VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 通过融合视觉与3D几何信息构建场景表征
  • 能区分细微但关键的动作差异,减少失败率
  • 适合需要高可靠性的机器人操作任务

视觉-语言-动作(VLA)模型在机器人操作中展现巨大潜力,但其测试时可靠性受限于单次动作预测,微小误差即可能导致抓取失败、碰撞或任务进展错误。一种自然替代方案是为VLA系统引入测试时验证,允许提出多个候选动作并评估后再执行。然而,可靠的动作验证极具挑战,因需区分候选动作间细微的几何差异,并判断动作是否真正推进任务目标。我们提出VeriSpace,一种用于VLA系统测试时动作选择的3D感知验证器。VeriSpace通过两个核心组件实现:双路径3D注入场景编码,联合保留视觉语义与显式3D几何;空间接地动作推理,通过分析任务相关空间关系、几何合理性及预期目标进展来评估每个动作。二者协同提升了对细微但结果关键动作的判别能力,且完全兼容现有VLA策略。在公开基准和真实机器人操作任务上的实验表明,VeriSpace在分布内与分布外设置下均显著提升决策可靠性,优于底层VLA策略及已有基于验证的方法。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have shown strong promise for robotic manipulation, but their reliability at test time remains limited by one-shot action prediction, where even small action errors can cause grasp failure, collision, or incorrect task progression. A natural alternative is to equip VLA systems with test-time verification, allowing multiple candidate actions to be proposed and evaluated before execution. However, reliable action verification is challenging because it requires not only distinguishing subtle geometric differences between candidate actions, but also assessing whether an action makes meaningful progress toward the task goal. We present VeriSpace, a 3D-aware action verifier for test-time action selection in VLA systems. VeriSpace evaluates candidate actions through two key components: Dual-Path 3D-Injected Scene Encoding, which constructs a scene representation that jointly preserves visual semantics and explicit 3D geometry, and Spatially-Grounded Action Reasoning, which evaluates each action by reasoning over task-relevant spatial relations, geometric validity, and expected goal progress. Together, these components enable more reliable discrimination between subtle yet outcome-critical action candidates while remaining fully compatible with existing VLA policies. Experiments on public benchmarks and real-world robotic manipulation tasks show that VeriSpace consistently improves decision reliability over both underlying VLA policies and prior verification-based methods, yielding substantial gains in both in-distribution and out-of-distribution settings.

机器人操作动作验证3D感知VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。