arXiv:2510.19818cs.LGcs.AI2025-10被引 7

用语义问答方式构建世界模型,提升机器人规划能力

Semantic World Models

  • 将世界模型建模为未来帧的语义问题问答任务
  • 在开放任务上实现显著优于传统像素重建方法的泛化性能
  • 可复用预训练视觉语言模型,适合需要强泛化的机器人决策场景

基于世界模型的规划为机器人控制提供强大范式。传统方法训练模型以当前帧和动作条件预测未来帧,但像素级重建目标常与实际规划目标不一致;强像素重建未必带来良好决策。本文提出:世界模型无需重建未来像素,只需预测任务相关的语义信息。为此,将世界建模转化为对未来帧中语义信息的视觉问答问题。这一视角使世界建模可采用视觉语言模型的核心工具。通过在图像-动作-文本数据上进行监督微调,可训练出“语义”世界模型,支持决策规划,并继承预训练视觉语言模型的泛化与鲁棒性。实验表明,该方法在开放式机器人任务中实现显著的泛化提升,优于典型的基于重建的动作条件世界模型。

原文摘要 · Abstract (English)

Planning with world models offers a powerful paradigm for robotic control. Conventional approaches train a model to predict future frames conditioned on current frames and actions, which can then be used for planning. However, the objective of predicting future pixels is often at odds with the actual planning objective; strong pixel reconstruction does not always correlate with good planning decisions. This paper posits that instead of reconstructing future frames as pixels, world models only need to predict task-relevant semantic information about the future. For such prediction the paper poses world modeling as a visual question answering problem about semantic information in future frames. This perspective allows world modeling to be approached with the same tools underlying vision language models. Thus vision language models can be trained as "semantic" world models through a supervised finetuning process on image-action-text data, enabling planning for decision-making while inheriting many of the generalization and robustness properties from the pretrained vision-language models. The paper demonstrates how such a semantic world model can be used for policy improvement on open-ended robotics tasks, leading to significant generalization improvements over typical paradigms of reconstruction-based action-conditional world modeling. Website available at https://weirdlabuw.github.io/swm.

世界模型机器人控制视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。