arXiv:2505.06535cs.AIcs.LG2025-05NeurIPS被引 1

用扩散模型动态找目标,省钱省力还透明。

Online Feedback Efficient Active Target Discovery in Partially Observable Environments

  • 基于扩散机制构建环境信念分布,自动权衡探索与利用。
  • 在有限采样次数内发现目标,性能超越多数基线方法。
  • 无需监督训练,适合医疗、遥感等数据昂贵场景。

在医学成像、环境监测或遥感等数据获取成本高昂的领域,如何在有限采样预算下高效发现目标至关重要。本文提出扩散引导的主动目标发现方法(DiffATD),通过维护环境中每个未观测状态的信念分布,动态平衡探索与利用。探索通过选择预期熵最高的区域降低不确定性,利用则聚焦于信念分布中目标可能性最高的区域,结合增量训练的奖励模型学习目标特征。DiffATD在部分可观测环境中实现高效目标发现,且无需任何监督训练。实验覆盖医学影像、物种发现和遥感等多个领域,结果表明其性能显著优于基线方法,并可与在全可观测条件下运行的监督方法媲美。该方法具备良好可解释性,区别于依赖大量标注数据的黑箱策略。

原文摘要 · Abstract (English)

In various scientific and engineering domains, where data acquisition is costly--such as in medical imaging, environmental monitoring, or remote sensing--strategic sampling from unobserved regions, guided by prior observations, is essential to maximize target discovery within a limited sampling budget. In this work, we introduce Diffusion-guided Active Target Discovery (DiffATD), a novel method that leverages diffusion dynamics for active target discovery. DiffATD maintains a belief distribution over each unobserved state in the environment, using this distribution to dynamically balance exploration-exploitation. Exploration reduces uncertainty by sampling regions with the highest expected entropy, while exploitation targets areas with the highest likelihood of discovering the target, indicated by the belief distribution and an incrementally trained reward model designed to learn the characteristics of the target. DiffATD enables efficient target discovery in a partially observable environment within a fixed sampling budget, all without relying on any prior supervised training. Furthermore, DiffATD offers interpretability, unlike existing black--box policies that require extensive supervised training. Through extensive experiments and ablation studies across diverse domains, including medical imaging, species discovery, and remote sensing, we show that DiffATD performs significantly better than baselines and competitively with supervised methods that operate under full environmental observability.

目标发现扩散模型主动学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。