让机器人精准按语言指令放物体,靠视觉目标引导实现零样本泛化。
AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement

- 先将语言转为视觉标记,再用目标引导策略执行放置。
- 在九类任务上零样本表现超越基线模型,精度达厘米级。
- 专设仿真基准SlotBench,评估复杂空间推理能力。
视觉-语言-动作(VLA)策略已成为通用机器人操作的有力范式。然而,面对组合性语言指令时,端到端的VLA策略在精确物体放置方面仍存在挑战。槽位级放置需要可靠的槽位定位和厘米级几何精度。为此,我们提出AnySlot框架,通过引入语言理解与控制之间的显式空间视觉目标,降低组合复杂度。AnySlot将语言指令转化为在目标槽位处渲染的空间标记,再由目标条件化的VLA策略执行该目标。这种分层设计将高层槽位选择与低层执行解耦,提升了语义准确性和空间鲁棒性。此外,针对此类高精度任务缺乏评测基准的问题,我们构建了SlotBench——一个包含九个任务类别的结构化仿真基准,用于评估槽位级放置中的空间推理能力。大量实验表明,AnySlot在零样本槽位级放置任务中显著优于平面式VLA基线和模块化定位方法。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) policies have emerged as a versatile paradigm for generalist robotic manipulation. However, precise object placement under compositional language remains challenging for end-to-end VLA policies. Slot-level placement requires reliable slot grounding and centimeter-level geometric precision. To this end, we propose AnySlot, a framework that reduces compositional complexity by introducing an explicit spatial visual goal between language grounding and control. AnySlot converts language into a visual goal by rendering a spatial marker at the intended slot, then executes this goal with a goal-conditioned VLA policy. This hierarchical design decouples high-level slot selection from low-level execution, improving semantic accuracy and spatial robustness. Furthermore, recognizing the lack of benchmarks for such precision-demanding tasks, we introduce SlotBench, a structured simulation benchmark with nine task categories for evaluating spatial reasoning in slot-level placement. Extensive experiments show that AnySlot significantly outperforms flat VLA baselines and modular grounding methods in zero-shot slot-level placement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。