用手绘草图控制风景动态照片生成,让创意更直观。
Sketch-Guided Stylized Landscape Cinemagraph Synthesis
- 用草图+文本联合控制生成风格化动态风景图
- 支持流体区域运动场精准建模,实现连续时间流动
- 适合艺术设计、数字媒体创作人群使用
由于难以定制复杂且富有表现力的动态元素,设计风格化动态照片颇具挑战。为实现对生成结果的直观与精细控制,手绘草图提供了一种超越文本输入的个性化设计表达方式。本文提出Sketch2Cinemagraph,一种基于草图引导的框架,可从自由手绘草图中条件生成风格化动态照片。该框架首先利用文本提示生成目标风格化景观图像及其真实版本,再通过预训练的目标检测模型获取流体区域掩码。随后,提出一个潜在运动扩散模型,在生成的景观图像流体区域内估计运动场,并以输入的运动草图为条件,结合提示控制掩码区域内生成的运动场。最后,基于U-Net的帧生成器在每个时间步将流体区域像素扭曲至目标位置,合成动态照片帧。实验验证了Sketch2Cinemagraph能从草图输入生成具有连续时序流动的美学上令人满意的风格化动态照片。我们通过定性与定量对比,展示了其相较于当前最优方法的优势。
原文摘要 · Abstract (English)
Designing stylized cinemagraphs is challenging due to the difficulty in customizing complex and expressive flow elements. To achieve intuitive and detailed control of the generated cinemagraphs, sketches provide a feasible solution to convey personalized design requirements beyond text inputs. In this paper, we propose Sketch2Cinemagraph, a sketch-guided framework that enables the conditional generation of stylized cinemagraphs from freehand sketches. Sketch2Cinemagraph adopts text prompts for initial landscape generation and provides sketch controls for both spatial and motion cues. The latent diffusion model first generates target stylized landscape images along with realistic versions. Then, a pre-trained object detection model obtains masks for the flow regions. We propose a latent motion diffusion model to estimate motion field in fluid regions of the generated landscape images. The input motion sketches serve as the conditions to control the generated motion fields in the masked fluid regions with the prompt. To synthesize cinemagraph frames, the pixels within fluid regions are warped to target locations at each timestep using a U-Net based frame generator. The results verified that Sketch2Cinemagraph can generate aesthetically appealing stylized cinemagraphs with continuous temporal flow from sketch inputs. We showcase the advantages of Sketch2Cinemagraph through qualitative and quantitative comparisons against the state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。