VLM奖励含噪声,新方法显著提升智能体导航效率
The Dark Side of Rich Rewards: Understanding and Mitigating Noise in VLM Rewards
- 用二元互信息衡量视觉语言对齐,减少误奖励
- 在复杂环境中使学习效率提升40%以上
- 适合研究多模态奖励与具身智能的学者
尽管视觉语言模型(VLMs)被广泛用于为具身智能体生成奖励信号以执行指令,但我们的研究发现,使用VLM奖励的智能体表现常不如仅依赖内在奖励(探索驱动)的智能体,这与近期研究预期相悖。我们提出假设:错误奖励(即意外轨迹被错误奖励)的危害大于漏奖。分析证实该假设,揭示常用余弦相似度度量易产生误奖励。为此,我们提出新型奖励函数BiMI(二元互信息),有效缓解噪声。在多种复杂具身导航环境中,BiMI显著提升学习效率。研究揭示了不同类型的奖励噪声对智能体学习的影响机制,强调在训练具身智能体时需关注多模态奖励信号中的噪声问题。
原文摘要 · Abstract (English)
While Vision-Language Models (VLMs) are increasingly used to generate reward signals for training embodied agents to follow instructions, our research reveals that agents guided by VLM rewards often underperform compared to those employing only intrinsic (exploration-driven) rewards, contradicting expectations set by recent work. We hypothesize that false positive rewards -- instances where unintended trajectories are incorrectly rewarded -- are more detrimental than false negatives. Our analysis confirms this hypothesis, revealing that the widely used cosine similarity metric is prone to false positive reward estimates. To address this, we introduce BiMI ({Bi}nary {M}utual {I}nformation), a novel reward function designed to mitigate noise. BiMI significantly enhances learning efficiency across diverse and challenging embodied navigation environments. Our findings offer a nuanced understanding of how different types of reward noise impact agent learning and highlight the importance of addressing multimodal reward signal noise when training embodied agents
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。