arXiv:2605.15519cs.CVcs.AI2026-05中稿 · AAMAS 2026被引 1

用扩散模型重建局部视野,实现多目标实时搜寻

DiffVAS: Diffusion-Guided Visual Active Search in Partially Observable Environments

论文配图:DiffVAS: Diffusion-Guided Visual Active Search in Partially Observable Environments
图 1 · 摘自论文原文
  • 通过扩散模型从零散视角重建完整地理区域
  • 在部分可观测环境下同时搜索多种目标,性能显著超越现有方法
  • 适合无人机巡检、搜救等真实场景的多目标搜索任务

视觉主动搜索(VAS)是一种利用视觉线索指导空中平台(如无人机)探索并定位大范围地理区域中兴趣区域的建模框架。潜在应用包括发现稀有野生动物盗猎热点、协助搜救任务及发现非法武器贩运等。以往方法假设整个搜索空间事先已知,这在受限视场和高采集成本下不现实,且通常针对特定目标训练策略,难以同时搜索多个类别。本文提出DiffVAS,一种在部分可观测环境中根据任务需求同时搜索多样目标的条件化策略,推动了视觉主动搜索在真实场景中的部署。DiffVAS利用扩散模型从连续观测的部分图像中重建完整地理区域,使基于强化学习的规划模块能够有效推理并指导后续搜索步骤。大量实验表明,DiffVAS在多个数据集上显著优于当前最优方法,在部分可观测环境下对多种目标的搜索表现突出。

原文摘要 · Abstract (English)

Visual active search (VAS) has been introduced as a modeling framework that leverages visual cues to direct aerial (e.g., UAV-based) exploration and pinpoint areas of interest within extensive geospatial regions. Potential applications of VAS include detecting hotspots for rare wildlife poaching, aiding search-and-rescue missions, and uncovering illegal trafficking of weapons, among other uses. Previous VAS approaches assume that the entire search space is known upfront, which is often unrealistic due to constraints such as a restricted field of view and high acquisition costs, and they typically learn policies tailored to specific target objects, which limits their ability to search for multiple target categories simultaneously. In this work, we propose DiffVAS, a target-conditioned policy that searches for diverse objects simultaneously according to task requirements in partially observable environments, which advances the deployment of visual active search policies in real-world applications. DiffVAS leverages a diffusion model to reconstruct the entire geospatial area from sequentially observed partial glimpses, which enables a target-conditioned reinforcement learning-based planning module to effectively reason and guide subsequent search steps. Extensive experiments demonstrate that DiffVAS excels in searching diverse objects in partially observable environments, significantly surpassing state-of-the-art methods on several datasets.

视觉搜索扩散模型无人机强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。