arXiv:2604.13793cs.CV2026-04被引 1

通过插值消除视角同步带来的跳跃,实现更自然的第一视角视频生成。

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation

论文配图:From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
图 1 · 摘自论文原文
  • 将第三人称到第一人称视频生成重构为连续序列建模问题。
  • 仅插值视频即可显著提升生成质量,证明时空不连续是主要障碍。
  • 框架通用性强,可统一处理跨视角视频生成任务,适合视频合成研究者。

第三人称到第一人称视频生成旨在从同步的第三人称视频和对应相机位姿中合成第一人称视频。尽管存在成对监督信号,但同步数据本身引入了显著的时空与几何不连续性,违背了标准视频生成模型对平滑运动的假设。本文识别出这种同步引起的跳跃为关键挑战,提出 Syn2Seq-Forcing 方法,通过在源视频与目标视频间进行插值,构建单一连续信号。通过将 Exo2Ego 重构为序列信号建模而非传统条件-输出任务,该方法使基于扩散的序列模型(如 Diffusion Forcing Transformers, DFoT)能更有效地捕捉帧间连贯过渡。实验表明,仅插值视频而不进行位姿插值即可带来显著性能提升,凸显时空不连续性的主导作用。此外,该框架具备通用性和灵活性,可统一建模 Exo2Ego 与 Ego2Exo 生成,为跨视角视频合成研究提供原则性基础。

原文摘要 · Abstract (English)

Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired supervision is available, synchronized exo-ego data inherently introduces substantial spatio-temporal and geometric discontinuities, violating the smooth-motion assumptions of standard video generation benchmarks. We identify this synchronization-induced jump as the central challenge and propose Syn2Seq-Forcing, a sequential formulation that interpolates between the source and target videos to form a single continuous signal. By reframing Exo2Ego as sequential signal modeling rather than a conventional condition-output task, our approach enables diffusion-based sequence models, e.g. Diffusion Forcing Transformers (DFoT), to capture coherent transitions across frames more effectively. Empirically, we show that interpolating only the videos, without performing pose interpolation already produces significant improvements, emphasizing that the dominant difficulty arises from spatio-temporal discontinuities. Beyond immediate performance gains, this formulation establishes a general and flexible framework capable of unifying both Exo2Ego and Ego2Exo generation within a single continuous sequence model, providing a principled foundation for future research in cross-view video synthesis.

视频生成扩散模型跨视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。