arXiv:2505.11350cs.RO2025-05中稿 · CoRL被引 2

用卫星图提升机器人野外视觉搜索效率,动态修正错误预测。

Search-TTA: A Multimodal Test-Time Adaptation Framework for Visual Search in the Wild

  • 结合卫星图与CLIP,用不确定性加权更新改进目标定位。
  • 在38万张图像数据集上,搜索效率最高提升30%。
  • 适配多种传感器和规划方法,支持真实无人机部署。

为实现户外视觉导航与搜索,机器人可利用卫星影像生成视觉先验,辅助高层搜索策略,即使这些影像分辨率不足以识别目标。然而,现有方法或假设无先验信息,或未考虑先验来源。近期工作使用大规模视觉语言模型(VLMs)生成通用先验,但其输出易因幻觉导致不准确,造成搜索效率低下。为此,我们提出Search-TTA,一种多模态测试时自适应框架,兼容多种输入模态(如图像、文本、声音)和规划方法(如基于强化学习)。首先,预训练卫星图像编码器,使其与CLIP视觉编码器对齐,输出目标存在概率分布用于视觉搜索;其次,我们的TTA框架在搜索过程中动态优化CLIP预测,采用受空间泊松点过程启发的不确定性加权梯度更新。为训练与评估Search-TTA,我们构建了基于互联网规模生态数据的AVS-Bench视觉搜索数据集,包含38万张图像与分类数据。实验表明,Search-TTA在初始CLIP预测较差(因领域差异或训练数据少)时,可使规划器性能提升最高达30.0%。其表现媲美更大规模的VLMs,且通过涌现对齐实现零样本跨模态泛化。最后,我们在真实无人机上通过软硬件联合测试验证,模拟其在大型仿真环境中执行任务。

原文摘要 · Abstract (English)

To perform outdoor visual navigation and search, a robot may leverage satellite imagery to generate visual priors. This can help inform high-level search strategies, even when such images lack sufficient resolution for target recognition. However, many existing informative path planning or search-based approaches either assume no prior information, or use priors without accounting for how they were obtained. Recent work instead utilizes large Vision Language Models (VLMs) for generalizable priors, but their outputs can be inaccurate due to hallucination, leading to inefficient search. To address these challenges, we introduce Search-TTA, a multimodal test-time adaptation framework with a flexible plug-and-play interface compatible with various input modalities (e.g., image, text, sound) and planning methods (e.g., RL-based). First, we pretrain a satellite image encoder to align with CLIP's visual encoder to output probability distributions of target presence used for visual search. Second, our TTA framework dynamically refines CLIP's predictions during search using uncertainty-weighted gradient updates inspired by Spatial Poisson Point Processes. To train and evaluate Search-TTA, we curate AVS-Bench, a visual search dataset based on internet-scale ecological data containing 380k images and taxonomy data. We find that Search-TTA improves planner performance by up to 30.0%, particularly in cases with poor initial CLIP predictions due to domain mismatch and limited training data. It also performs comparably with significantly larger VLMs, and achieves zero-shot generalization via emergent alignment to unseen modalities. Finally, we deploy Search-TTA on a real UAV via hardware-in-the-loop testing, by simulating its operation within a large-scale simulation that provides onboard sensing.

视觉搜索测试时自适应多模态无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。