用文字训练视觉决策模型,减少对图像数据依赖。
Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
- 在文本描述场景中用强化学习训练模型推理能力
- 跨任务测试中表现优于传统微调方法,通用性强
- 适合需要少样本视觉决策的应用场景
视觉语言模型在多项任务中表现优异,但在复杂决策中常缺乏情境推理能力。本文发现,当视觉场景被文本描述替代时,VLM可实现出色的决策性能,表明基础推理能力可通过语言有效学习。基于此,我们提出Praxis-VLM,一种用于视觉接地决策的推理型VLM。该模型采用GRPO算法,在文本场景中训练模型评估动作及其后果,从而获得强推理能力。这些纯文本习得的推理技能能有效迁移到多模态推理中,显著降低对稀缺图像-文本配对数据的依赖。在多个决策基准上的实验表明,Praxis-VLM明显优于标准监督微调,展现出更强性能与泛化能力。进一步分析证实,模型进行显式且高效的推理,支撑其优越表现与适应性。
原文摘要 · Abstract (English)
Vision Language Models exhibit impressive performance for various tasks, yet they often lack the sophisticated situational reasoning required for complex decision-making. This paper shows that VLMs can achieve surprisingly strong decision-making performance when visual scenes are replaced by textual descriptions, suggesting foundational reasoning can be effectively learned from language. Motivated by this insight, we propose Praxis-VLM, a reasoning VLM for vision-grounded decision-making. Praxis-VLM employs the GRPO algorithm on textual scenarios to instill robust reasoning capabilities, where models learn to evaluate actions and their consequences. These reasoning skills, acquired purely from text, successfully transfer to multimodal inference with visual inputs, significantly reducing reliance on scarce paired image-text training data. Experiments across diverse decision-making benchmarks demonstrate that Praxis-VLM substantially outperforms standard supervised fine-tuning, exhibiting superior performance and generalizability. Further analysis confirms that our models engage in explicit and effective reasoning, underpinning their enhanced performance and adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。