arXiv:2508.21102cs.CVcs.RO2025-08中稿 · CoRL被引 1

提出GENNAV模型,精准定位多类导航区域的模糊边界。

GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions

  • 基于存在性预测与掩码生成联合建模,支持多目标分割
  • 在GRiN-Drive数据集上超越基线方法,尤其在多目标场景表现优异
  • 真实城市环境中零样本迁移能力强,适用于自动驾驶导航

本文聚焦于从自然语言指令和移动设备前视图像中识别目标区域位置的任务,该任务因需同时完成存在性判断与分割而具有挑战性,尤其对边界模糊的stuff类型目标区域。现有方法在处理此类目标时性能不足,且难以应对无目标或多个目标的情况。为此,我们提出GENNAV模型,可同时预测目标是否存在,并生成多个stuff类型目标区域的分割掩码。为评估该模型,我们构建了新基准GRiN-Drive,包含无目标、单目标和多目标三类样本。GENNAV在标准评估指标上优于基线方法。此外,我们在五个地理分布不同的城市中,使用四辆汽车进行了真实世界实验,验证其零样本迁移能力。结果表明,GENNAV在复杂真实环境中表现稳健,显著优于基线方法。

原文摘要 · Abstract (English)

We focus on the task of identifying the location of target regions from a natural language instruction and a front camera image captured by a mobility. This task is challenging because it requires both existence prediction and segmentation, particularly for stuff-type target regions with ambiguous boundaries. Existing methods often underperform in handling stuff-type target regions, in addition to absent or multiple targets. To overcome these limitations, we propose GENNAV, which predicts target existence and generates segmentation masks for multiple stuff-type target regions. To evaluate GENNAV, we constructed a novel benchmark called GRiN-Drive, which includes three distinct types of samples: no-target, single-target, and multi-target. GENNAV achieved superior performance over baseline methods on standard evaluation metrics. Furthermore, we conducted real-world experiments with four automobiles operated in five geographically distinct urban areas to validate its zero-shot transfer performance. In these experiments, GENNAV outperformed baseline methods and demonstrated its robustness across diverse real-world environments. The project page is available at https://gennav.vercel.app/.

导航定位语义分割自动驾驶多目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。