arXiv:2410.14250cs.CV2024-10NeurIPS被引 33

用能量模型让导航更贴近专家行为,减少错误累积。

Vision-Language Navigation with Energy-Based Policy

  • 用能量模型建模状态-动作联合分布,低能量对应专家高频动作
  • 在多个数据集上显著提升导航准确率,超越传统监督学习方法
  • 适合希望改进视觉语言导航泛化能力的研究者

视觉语言导航(VLN)要求智能体根据人类指令执行动作。现有模型通常通过专家示范进行监督学习或手动设计奖励函数优化,但这些方法忽略了马尔可夫决策过程中的误差累积问题,难以匹配专家策略的分布。为此,我们提出能量基础导航策略(ENP),利用能量模型建模状态-动作联合分布:低能量值对应专家最可能采取的状态-动作对。理论上,该方法等价于最小化专家与自身占据测度之间的前向散度。因此,ENP通过最大化动作可能性并协同建模导航状态动态,实现与专家策略的全局对齐。在R2R、REVERIE、RxR和R2R-CE等多个基准上,结合不同架构的ENP均取得优异表现,充分释放了现有VLN模型潜力。

原文摘要 · Abstract (English)

Vision-language navigation (VLN) requires an agent to execute actions following human instructions. Existing VLN models are optimized through expert demonstrations by supervised behavioural cloning or incorporating manual reward engineering. While straightforward, these efforts overlook the accumulation of errors in the Markov decision process, and struggle to match the distribution of the expert policy. Going beyond this, we propose an Energy-based Navigation Policy (ENP) to model the joint state-action distribution using an energy-based model. At each step, low energy values correspond to the state-action pairs that the expert is most likely to perform, and vice versa. Theoretically, the optimization objective is equivalent to minimizing the forward divergence between the occupancy measure of the expert and ours. Consequently, ENP learns to globally align with the expert policy by maximizing the likelihood of the actions and modeling the dynamics of the navigation states in a collaborative manner. With a variety of VLN architectures, ENP achieves promising performances on R2R, REVERIE, RxR, and R2R-CE, unleashing the power of existing VLN models.

视觉语言导航能量模型策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。