arXiv:2606.14879cs.ROcs.CV2026-06

用视觉好奇心引导扩散策略,实现单目相机下的高效无地图探索

VANDERER: Map-Free Exploration using Future-Aware and Visual-Curiosity-Guided Diffusion Policy

论文配图:VANDERER: Map-Free Exploration using Future-Aware and Visual-Curiosity-Guided Diffusion Policy
图 1 · 摘自论文原文
  • 基于视觉好奇心模块预测动作后果,驱动扩散模型生成探索动作
  • 在模拟环境中平均比NoMaD多探索13.4%的区域
  • 适合传感器受限的机器人自主探索场景

移动智能体需高效探索策略以映射未知环境并自主规划任务。传统方法依赖构建占据图并优化未探索区域的访问顺序,但在仅使用单目相机等传感受限场景下,生成准确占据图极具挑战。为此,我们提出VANDERER,一种利用视觉好奇心模块(VCM)仅通过单目图像数据引导预训练扩散策略的探索框架。该模块通过导航世界模型预测动作结果,并以好奇心成本评估其价值,进而引导扩散过程生成最大化探索的动作。在多样化模拟环境中评估显示,VANDERER始终优于现有基线,平均比NoMaD多探索13.4%的区域。结果表明,在户外环境中视觉好奇心与几何好奇心存在直接相关性,证明VANDERER能有效利用此关系,实现传感器受限智能体的高效探索。

原文摘要 · Abstract (English)

Mobile agents require efficient exploration strategies to map unseen environments and autonomously plan tasks. Traditional methods rely on generating occupancy maps and optimizing the sequence in which unexplored regions are visited. However, in sensor-constrained settings, such as those limited to monocular cameras, generating accurate occupancy maps is challenging. To address this, we propose VANDERER, an exploration framework that leverages a Visual Curiosity Module (VCM) to guide pre-trained diffusion policies using only monocular image data. This curiosity module predicts the outcomes of proposed actions via a navigation world model and evaluates them through a curiosity cost. The cost then guides the diffusion process toward generating actions that maximize exploration. Evaluated across diverse simulated environments, VANDERER consistently outperforms established baselines, exploring an average of 13.4% more area than NoMaD. Our results reveal a direct correlation between visual and geometric curiosity in outdoor environments, demonstrating that VANDERER can effectively leverage this relationship for efficient exploration using sensor-constrained agents.

无地图探索扩散模型视觉好奇心

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。