arXiv:2511.06240cs.ROcs.AI2025-11AAAI被引 3

用视觉语言模型指导机器人选址,提升复杂任务成功率

Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation

  • 融合视觉语言模型与几何约束,分阶段优化机器人基座位置
  • 在5个任务中达85%成功率,显著优于传统方法
  • 适合需要理解指令并自主选址的移动操作场景

在开放词汇移动操作(OVMM)中,任务成功常取决于机器人的基座位置选择。现有方法多基于距离接近性导航,忽视可用性特征,导致操作失败频发。本文提出一种零样本的渐进式基座选址框架,将视觉语言模型(VLMs)的语义理解与几何可行性通过迭代优化结合。构建跨模态表示:可用性RGB与障碍物图+,实现语义与空间上下文对齐,突破单视角RGB感知局限。利用VLM提供的粗粒度语义先验引导搜索至任务相关区域,并以几何约束精细调整位置,降低陷入局部最优风险。在五个多样化开放词汇移动操作任务上评估,系统达到85%成功率,显著优于经典几何规划器和VLM-based方法。结果表明,具备可用性感知与多模态推理能力的规划系统,在指令驱动的通用化移动操作中具有巨大潜力。

原文摘要 · Abstract (English)

In open-vocabulary mobile manipulation (OVMM), task success often hinges on the selection of an appropriate base placement for the robot. Existing approaches typically navigate to proximity-based regions without considering affordances, resulting in frequent manipulation failures. We propose Affordance-Guided Coarse-to-Fine Exploration, a zero-shot framework for base placement that integrates semantic understanding from vision-language models (VLMs) with geometric feasibility through an iterative optimization process. Our method constructs cross-modal representations, namely Affordance RGB and Obstacle Map+, to align semantics with spatial context. This enables reasoning that extends beyond the egocentric limitations of RGB perception. To ensure interaction is guided by task-relevant affordances, we leverage coarse semantic priors from VLMs to guide the search toward task-relevant regions and refine placements with geometric constraints, thereby reducing the risk of convergence to local optima. Evaluated on five diverse open-vocabulary mobile manipulation tasks, our system achieves an 85% success rate, significantly outperforming classical geometric planners and VLM-based methods. This demonstrates the promise of affordance-aware and multimodal reasoning for generalizable, instruction-conditioned planning in OVMM.

移动操作视觉语言模型基座选址多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。