让模型学会感知和操控视频时间流速,实现变速生成与超分辨率。
Seeing Fast and Slow: Learning the Flow of Time in Videos

- 自监督学习视频中的时间线索,判断快慢变化并估计播放速度。
- 构建迄今最大的慢动作视频数据集,提升时序细节表现力。
- 可实现指定速度生成视频与低帧率视频的高帧率重建。
如何判断视频是加速还是减速?如何生成不同速度的视频?尽管视频在现代计算机视觉中占据核心地位,但对时间流逝的感知与控制仍鲜有研究。本文将时间视为可学习的视觉概念,提出模型来推理和操控视频中的时间流。首先,利用视频中的多模态线索和时间结构,以自监督方式学习检测速度变化并估计播放速度;进而基于此,从嘈杂的真实场景数据中构建迄今最大的慢动作视频数据集——这类由高速摄像机拍摄的视频包含远超标准视频的时序细节。利用该数据,进一步开发出具备时间控制能力的模型,包括速度条件下的视频生成(按指定速度生成运动)和时间超分辨率(将低帧率模糊视频转化为高帧率、细粒度时序序列)。研究揭示时间是可操作的感知维度,为可控视频生成、时间伪造检测及更丰富的世界模型开辟新路径。
原文摘要 · Abstract (English)
How can we tell whether a video has been sped up or slowed down? How can we generate videos at different speeds? Although videos have been central to modern computer vision research, little attention has been paid to perceiving and controlling the passage of time. In this paper, we study time as a learnable visual concept and develop models for reasoning about and manipulating the flow of time in videos. We first exploit the multimodal cues and temporal structure naturally present in videos to learn, in a self-supervised manner, to detect speed changes and estimate playback speed. We then show that these learned temporal reasoning models enable us to curate the largest slow-motion video dataset to date from noisy in-the-wild sources. Such slow-motion footage, typically filmed by high-speed cameras, contains substantially richer temporal detail than standard videos. Using this data, we further develop models capable of temporal control, including speed-conditioned video generation, which produces motion at specified playback speed, and temporal super-resolution, which tranforms low-FPS, blurry videos into high-FPS sequences with fine-grained temporal details. Our findings highlight time as a manipulable, perceptual dimension in video learning, opening doors to temporally controllable video generation, temporal forensics detection, and potentially richer world-models that understand how events unfold over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。