arXiv:2511.19033cs.CV2025-11被引 3

让视觉语言模型更聪明地探索新环境,无需额外训练。

ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay

  • 推理时回放过往经验,用抽象知识增强决策
  • 在多个基准上成功率和导航效率提升达3倍
  • 适合希望不训练就能提升智能体探索能力的研究者

具身探索是目标驱动的过程,要求智能体具备精细感知与知识增强的决策能力。尽管近期研究利用多模态大模型(MLLM)因其强大的感知与推理能力进行探索,但现有方法仍存在三方面问题:(i) 依赖深度但过时的预训练知识;(ii) 仿射学习或强化学习等训练方法在长时程、稀疏奖励任务中成本高昂;(iii) 前沿探索产生庞大且视觉复杂的动作空间,MLLM难以可靠决策。为此,我们提出ReEXplore,一种无需训练的框架,通过推理时回放抽象经验注入知识,并采用分层前沿选择将前沿排序分解为粗粒度到细粒度决策。该方法实现稳健、可追溯、高效的探索。在多个具身探索基准测试中,相较于强基线,使用开源模型骨干时,性能提升最高达3倍,成功率达3倍,导航效率也显著提高。

原文摘要 · Abstract (English)

Embodied exploration is a target-driven process that requires embodied agents to possess fine-grained perception and knowledge-enhanced decision making. While recent attempts leverage MLLMs for exploration due to their strong perceptual and reasoning abilities, we find that MLLM-based embodied agents remain suboptimal in exploring new environments: (i) they rely on profound but stale pre-trained knowledge, (ii) training-based approaches such as imitation learning or reinforcement learning are expensive for long-horizon tasks with sparse outcome rewards, and (iii) frontier-based exploration yields a large, visually nuanced action space that is difficult for MLLMs to make reliable decisions. We address these challenges with ReEXplore, a training-free framework that performs retrospective experience replay to inject distilled, abstract experience at inference time, and hierarchical frontier selection to decompose frontier ranking into coarse-to-fine decisions. Our approach enables robust, traceable, and efficient exploration. Across multiple embodied exploration benchmarks, ReEXplore yields great improvements over strong MLLM baselines, up to 3x higher performance in both success rate and in navigation efficiency under open-source backbones.

具身智能多模态模型探索算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。