arXiv:2601.10781cs.CV2026-01被引 5

用语言引导的流模型预测未来光流,提升机器人控制与视频生成效果

Future Optical Flow Prediction Improves Robot Control & Video Generation

  • 融合视觉语言模型与扩散架构,实现多模态推理与像素级生成
  • 在海量无结构网络视频数据上训练,有效提取噪声中的运动信号
  • 跨控制与生成任务验证通用性,适合需要未来运动预测的场景

未来运动表征(如光流)在控制与生成任务中具有巨大价值。然而,如何泛化地预测空间密集的运动表征仍是关键挑战,且从真实世界噪声数据中学习此类预测仍相对未被探索。我们提出 FOFPred,一种新型语言引导的光流预测模型,结合统一的视觉-语言模型(VLM)与扩散架构。该组合支持强大的多模态推理与像素级生成保真度。模型在大规模网络人类活动数据上训练——这是一种高度可扩展但非结构化的数据源。为从这类噪声视频-文本对中提取有意义信号,我们采用关键的数据预处理技术及具备强大图像预训练能力的统一架构。训练后的模型被拓展至机器人操控与视频生成两个下游任务。在语言驱动设置下,于机器人操作与视频生成任务上的评估证实了 FOFPred 的跨领域通用性,验证了统一的 VLM-扩散架构与从多样化网络数据中可扩展学习对未来光流预测的价值。

原文摘要 · Abstract (English)

Future motion representations, such as optical flow, offer immense value for control and generative tasks. However, forecasting generalizable spatially dense motion representations remains a key challenge, and learning such forecasting from noisy, real-world data remains relatively unexplored. We introduce FOFPred, a novel language-conditioned optical flow forecasting model featuring a unified Vision-Language Model (VLM) and Diffusion architecture. This unique combination enables strong multimodal reasoning with pixel-level generative fidelity for future motion prediction. Our model is trained on web-scale human activity data-a highly scalable but unstructured source. To extract meaningful signals from this noisy video-caption data, we employ crucial data preprocessing techniques and our unified architecture with strong image pretraining. The resulting trained model is then extended to tackle two distinct downstream tasks in control and generation. Evaluations across robotic manipulation and video generation under language-driven settings establish the cross-domain versatility of FOFPred, confirming the value of a unified VLM-Diffusion architecture and scalable learning from diverse web data for future optical flow prediction.

光流预测扩散模型机器人控制视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。