让视觉语言模型通过互动学习物理常识,发现仍难泛化。
Can Vision Language Models Learn Intuitive Physics from Interaction?
- 用强化学习让模型在模拟环境里互动学习物理
- 互动训练提升任务内表现但无法推广到新任务
- 即使视觉和原理相同,模型也难以跨任务泛化
预训练的视觉语言模型缺乏对物理世界的直观理解。近期研究显示,监督微调可提升模型在简单物理任务上的表现,但微调后的模型似乎未能习得能泛化到新情境的稳健物理规则。基于认知科学的研究,我们假设模型需与环境互动才能真正学习其物理动态。为此,我们训练了通过与模拟环境互动学习的模型,采用强化学习方法。尽管互动学习提升了模型在特定任务内的表现,却未能生成具备可泛化物理直觉的模型。我们发现,即使任务共享相同的视觉统计特征和物理原理,从一个任务训练的模型也无法可靠地推广到相关任务,无论是否通过互动训练。
原文摘要 · Abstract (English)
Pre-trained vision language models do not have good intuitions about the physical world. Recent work has shown that supervised fine-tuning can improve model performance on simple physical tasks. However, fine-tuned models do not appear to learn robust physical rules that can generalize to new contexts. Based on research in cognitive science, we hypothesize that models need to interact with an environment to properly learn its physical dynamics. We train models that learn through interaction with a simulated environment using reinforcement learning. While learning from interaction allows models to improve their within-task performance, it fails to produce models with generalizable physical intuitions. We find that models trained on one task do not reliably generalize to related tasks, even if the tasks share visual statistics and physical principles, and regardless of whether the models are trained through interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。