arXiv:2601.03956cs.RO2026-01被引 1

让机器人学会主动搬东西清路,比传统方法成功率高17%。

CoINS: Counterfactual Interactive Navigation via Skill-Aware VLM

  • 用带技能理解的视觉语言模型预测搬哪件物品能通路
  • 在复杂场景中成功率达80%以上,比最好基线提升17%
  • 适合需要动手清理障碍的机器人导航任务

近期视觉语言模型(VLM)在机器人规划中展现出巨大潜力,但通常仅作为语义推理器,缺乏对机器人物理能力的内在理解。这一局限在交互式导航中尤为关键,因机器人需主动改造杂乱环境以创建可通行路径。现有基于VLM的导航系统多局限于被动避障,无法判断何时及如何与物体互动以清除阻塞路径。为此,我们提出基于技能感知的反事实交互导航框架CoINS,通过分层结构整合技能感知推理与稳健低层执行。具体而言,我们微调了一个名为InterNav-VLM的VLM,将技能可用性与具体约束参数引入输入上下文,并将其映射到度量尺度的环境表征中。通过在新提出的InterNav数据集上微调,模型内化了反事实推理逻辑,隐式评估物体移除对导航连通性的影响,从而决定互动必要性与目标选择。为执行生成的高层计划,我们通过强化学习构建了综合性技能库,特别引入面向通行性的策略以操作各类物体实现路径清理。我们在Isaac Sim中设计了系统性基准,用于评估交互导航的推理与执行性能。大量仿真与真实世界实验表明,CoINS显著优于代表性基线,在整体成功率上高出17%,在复杂长程场景中性能提升超过80%。

原文摘要 · Abstract (English)

Recent Vision-Language Models (VLMs) have demonstrated significant potential in robotic planning. However, they typically function as semantic reasoners, lacking an intrinsic understanding of the specific robot's physical capabilities. This limitation is particularly critical in interactive navigation, where robots must actively modify cluttered environments to create traversable paths. Existing VLM-based navigators are predominantly confined to passive obstacle avoidance, failing to reason about when and how to interact with objects to clear blocked paths. To bridge this gap, we propose Counterfactual Interactive Navigation via Skill-aware VLM (CoINS), a hierarchical framework that integrates skill-aware reasoning and robust low-level execution. Specifically, we fine-tune a VLM, named InterNav-VLM, which incorporates skill affordance and concrete constraint parameters into the input context and grounds them into a metric-scale environmental representation. By internalizing the logic of counterfactual reasoning through fine-tuning on the proposed InterNav dataset, the model learns to implicitly evaluate the causal effects of object removal on navigation connectivity, thereby determining interaction necessity and target selection. To execute the generated high-level plans, we develop a comprehensive skill library through reinforcement learning, specifically introducing traversability-oriented strategies to manipulate diverse objects for path clearance. A systematic benchmark in Isaac Sim is proposed to evaluate both the reasoning and execution aspects of interactive navigation. Extensive simulations and real-world experiments demonstrate that CoINS significantly outperforms representative baselines, achieving a 17\% higher overall success rate and over 80\% improvement in complex long-horizon scenarios compared to the best-performing baseline

交互导航视觉语言模型机器人控制技能学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。