arXiv:2602.07277cs.CVcs.LG2026-02被引 1

让智能体从多视角想象未来,提升规划能力。

Cross-View World Models

  • 用跨视角预测训练模型,学习环境3D结构的不变表示。
  • 在多视角数据上训练,使模型能跨视点准确预测未来状态。
  • 适合需要空间推理和多视角协作的任务场景。

世界模型使智能体能够通过预想未来状态来规划行为,但现有方法通常仅依赖单一视角(如第一人称视角),即使其他视角(如俯视图)更有利于规划。本文提出跨视角世界模型(XVWM),通过跨视角预测目标进行训练:给定一个视角下的帧序列,预测执行动作后在相同或不同视角下的未来状态。强制跨视角一致性作为几何正则化:由于输入与输出视角可能无视觉重叠,模型必须学习环境3D结构的视点不变表征。我们在Aimlabs平台获取的同步多视角游戏数据上训练,该平台提供高精度对齐的多相机录制与高频动作标签。训练后的模型支持多视角并行想象,使智能体在执行时仍以第一人称视角行动,却可在任意视角下进行规划。结果表明,多视角一致性为具身空间表征提供了强学习信号。此外,从他人视角预测自身行为后果,或可为多智能体环境中的换位思考奠定基础。

原文摘要 · Abstract (English)

World models enable agents to plan by imagining future states, but existing approaches operate from a single viewpoint, typically egocentric, even when other perspectives would make planning easier; navigation, for instance, benefits from a bird's-eye view. We introduce Cross-View World Models (XVWM), trained with a cross-view prediction objective: given a sequence of frames from one viewpoint, predict the future state from the same or a different viewpoint after an action is taken. Enforcing cross-view consistency acts as geometric regularization: because the input and output views may share little or no visual overlap, to predict across viewpoints, the model must learn view-invariant representations of the environment's 3D structure. We train on synchronized multi-view gameplay data from Aimlabs, an aim-training platform providing precisely aligned multi-camera recordings with high-frequency action labels. The resulting model gives agents parallel imagination streams across viewpoints, enabling planning in whichever frame of reference best suits the task while executing from the egocentric view. Our results show that multi-view consistency provides a strong learning signal for spatially grounded representations. Finally, predicting the consequences of one's actions from another viewpoint may offer a foundation for perspective-taking in multi-agent settings.

世界模型多视角空间推理规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。