用外视角视频和文本生成第一人称视角视频,关键在模拟手物交互。
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
- 通过跨视角建模手物交互,预测未来第一人称画面。
- 在两个数据集上优于现有模型,手和物体生成更准确。
- 适合增强现实、智能体研究者关注,可生成真实交互内容。
第一人称视频生成在增强现实与具身智能领域具有广泛应用前景。本文探索跨视角视频预测任务:给定一个外视角视频、对应的第一人称视频首帧及文本指令,目标是生成后续第一人称视频帧。受第一人称视频中手物交互(HOI)体现主体意图的启发,我们提出 EgoExo-Gen,显式建模手物动态以实现跨视角视频预测。该方法包含两阶段:首先设计跨视角 HOI 掩码预测模型,通过建模时空上的内外视角对应关系,预测未来第一人称帧中的手物交互掩码;随后利用视频扩散模型,结合首帧和文本指令,以 HOI 掩码为结构引导生成未来帧。为支持训练,我们构建自动化流水线,借助视觉基础模型为内外视角视频生成伪 HOI 掩码。大量实验表明,EgoExo-Gen 在 Ego-Exo4D 与 H2O 基准数据集上均优于现有视频预测模型,且 HOI 掩码显著提升手部与交互物体的生成质量。
原文摘要 · Abstract (English)
Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and textual instructions, the goal is to generate futur frames of the ego-centric video. Inspired by the notion that hand-object interactions (HOI) in ego-centric videos represent the primary intentions and actions of the current actor, we present EgoExo-Gen that explicitly models the hand-object dynamics for cross-view video prediction. EgoExo-Gen consists of two stages. First, we design a cross-view HOI mask prediction model that anticipates the HOI masks in future ego-frames by modeling the spatio-temporal ego-exo correspondence. Next, we employ a video diffusion model to predict future ego-frames using the first ego-frame and textual instructions, while incorporating the HOI masks as structural guidance to enhance prediction quality. To facilitate training, we develop an automated pipeline to generate pseudo HOI masks for both ego- and exo-videos by exploiting vision foundation models. Extensive experiments demonstrate that our proposed EgoExo-Gen achieves better prediction performance compared to previous video prediction models on the Ego-Exo4D and H2O benchmark datasets, with the HOI masks significantly improving the generation of hands and interactive objects in the ego-centric videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。