用手绘草图直接生成高质量视频动画,支持不同绘画水平用户。
VidSketch: Hand-drawn Sketch-Driven Video Generation with Diffusion Control
- 通过分层草图控制策略动态调节草图引导强度,适应不同画技。
- 引入时空注意力机制,显著提升视频帧间一致性与流畅度。
- 首个支持多草图+文本生成视频的工具,适合创意设计初学者。
随着生成式人工智能的发展,已有研究能够从手绘草图生成美观静态图像,满足大众绘画需求。然而,现有方法局限于静态图像生成,无法利用手绘草图控制视频动画生成。为填补这一空白,我们提出VidSketch,是首个能直接从任意数量手绘草图和简单文本提示生成高质量视频动画的方法,拉近普通用户与专业艺术家之间的距离。具体而言,该方法提出一种分层草图控制策略,可自动调整生成过程中草图的引导强度,适应不同绘画水平的用户;此外,设计了时空注意力机制,有效增强生成视频在时空维度的一致性,显著提升帧间连贯性。更多案例可访问官网查看。
原文摘要 · Abstract (English)
With the advancement of generative artificial intelligence, previous studies have achieved the task of generating aesthetic images from hand-drawn sketches, fulfilling the public's needs for drawing. However, these methods are limited to static images and lack the ability to control video animation generation using hand-drawn sketches. To address this gap, we propose VidSketch, the first method capable of generating high-quality video animations directly from any number of hand-drawn sketches and simple text prompts, bridging the divide between ordinary users and professional artists. Specifically, our method introduces a Level-Based Sketch Control Strategy to automatically adjust the guidance strength of sketches during the generation process, accommodating users with varying drawing skills. Furthermore, a TempSpatial Attention mechanism is designed to enhance the spatiotemporal consistency of generated video animations, significantly improving the coherence across frames. You can find more detailed cases on our official website.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。