arXiv:2504.03153cs.LG2025-04被引 1

融合视觉与文本的强化学习框架,提升机器人实验室内决策效率

MORAL: A Multimodal Reinforcement Learning Framework for Decision Making in Autonomous Laboratories

  • 用图像+文本早期融合增强决策输入
  • 任务完成率提升20%,优于纯视觉/文本模型
  • 适合自动化实验室与具身智能系统研究者

我们提出MORAL(多模态强化学习框架),通过整合视觉与文本输入,提升自主机器人实验室中的序列决策能力。基于BridgeData V2数据集,使用预训练的BLIP-2模型生成细调后的图像描述,并与视觉特征进行早期融合。融合表示通过Deep Q-Network(DQN)和近端策略优化(PPO)代理处理。实验表明,多模态代理在充分训练后,任务完成率比基线提升20%,显著优于仅使用视觉或文本的模型。相比基于Transformer和循环结构的多模态强化学习模型,本方法在累积奖励和描述质量指标(BLEU、METEOR、ROUGE-L)上表现更优。结果表明语义对齐的语言提示能有效提升代理的学习效率与泛化能力。该框架推动了多模态强化学习及动态现实环境中具身智能系统的发展。

原文摘要 · Abstract (English)

We propose MORAL (a multimodal reinforcement learning framework for decision making in autonomous laboratories) that enhances sequential decision-making in autonomous robotic laboratories through the integration of visual and textual inputs. Using the BridgeData V2 dataset, we generate fine-tuned image captions with a pretrained BLIP-2 vision-language model and combine them with visual features through an early fusion strategy. The fused representations are processed using Deep Q-Network (DQN) and Proximal Policy Optimization (PPO) agents. Experimental results demonstrate that multimodal agents achieve a 20% improvement in task completion rates and significantly outperform visual-only and textual-only baselines after sufficient training. Compared to transformer-based and recurrent multimodal RL models, our approach achieves superior performance in cumulative reward and caption quality metrics (BLEU, METEOR, ROUGE-L). These results highlight the impact of semantically aligned language cues in enhancing agent learning efficiency and generalization. The proposed framework contributes to the advancement of multimodal reinforcement learning and embodied AI systems in dynamic, real-world environments.

多模态强化学习机器人实验具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。