arXiv:2601.02646cs.CVcs.AI2026-01被引 1

用一张照片生成可控制的循环动态影像,支持自由轨迹与时间调节。

DreamLoop: Controllable Cinemagraph Generation from a Single Photograph

  • 基于视频扩散模型,通过时序衔接与运动条件双重训练实现可控生成
  • 输入图像首尾一致强制无缝循环,静态背景保持不变
  • 用户指定运动路径即可精准控制目标物体的动画轨迹和节奏

Cinemagraphs 是将静态照片与局部循环运动结合的独特艺术形式。从单张照片中可控生成 cinemagraphs 仍具挑战性。现有图像动画技术仅限于水、烟等重复纹理的简单低频运动,而大规模视频扩散模型未针对 cinemagraph 约束设计,缺乏专用数据以生成无缝可控的循环内容。我们提出 DreamLoop,一个无需 cinemagraph 训练数据的可控视频合成框架。核心思想是通过时序衔接和运动条件两个目标对通用视频扩散模型进行微调,实现灵活生成。推理时,以输入图作为首尾帧,强制形成无缝循环;通过静态轨迹约束保持背景静止;用户指定目标对象的运动路径,即可直观控制动画轨迹与时间。据我们所知,DreamLoop 是首个可在一般场景中实现灵活、直观控制的 cinemagraph 生成方法,生成结果质量高且符合用户意图,优于现有方法。

原文摘要 · Abstract (English)

Cinemagraphs, which combine static photographs with selective, looping motion, offer unique artistic appeal. Generating them from a single photograph in a controllable manner is particularly challenging. Existing image-animation techniques are restricted to simple, low-frequency motions and operate only in narrow domains with repetitive textures like water and smoke. In contrast, large-scale video diffusion models are not tailored for cinemagraph constraints and lack the specialized data required to generate seamless, controlled loops. We present DreamLoop, a controllable video synthesis framework dedicated to generating cinemagraphs from a single photo without requiring any cinemagraph training data. Our key idea is to adapt a general video diffusion model by training it on two objectives: temporal bridging and motion conditioning. This strategy enables flexible cinemagraph generation. During inference, by using the input image as both the first- and last- frame condition, we enforce a seamless loop. By conditioning on static tracks, we maintain a static background. Finally, by providing a user-specified motion path for a target object, our method provides intuitive control over the animation's trajectory and timing. To our knowledge, DreamLoop is the first method to enable cinemagraph generation for general scenes with flexible and intuitive controls. We demonstrate that our method produces high-quality, complex cinemagraphs that align with user intent, outperforming existing approaches.

视频生成可控生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。