arXiv:2503.23736cs.CVcs.MM2025-03被引 2

无需训练即可将静态画作转化为动态视频,保持原画风格同时精准响应文本指令。

Every Painting Awakened: A Training-free Framework for Painting-to-Animation Generation

  • 利用预训练模型的图文对齐能力,通过合成图像引导动画生成。
  • 双路径评分蒸馏与混合潜在空间融合,实现动态与画风双重保真。
  • 零训练参数,可直接接入现有视频生成模型,适合艺术创作与数字修复。

我们提出一种无需训练的框架,专为将真实世界静态画作转化为动态视频而设计,解决在文本引导下运动与原作风格难以对齐的问题。现有图像到视频方法主要基于自然视频数据集训练,常无法从静态画作中生成合理动态内容,导致两类失效模式:一是因文本运动理解不足导致输出静态;二是因与真实艺术风格对齐不佳导致动态失真。本方法借助预训练图像模型强大的图文对齐能力,引入两种创新机制:(1)双路径评分蒸馏:采用双路径架构,从真实与合成数据中蒸馏运动先验,保留原画静态细节,学习合成帧的动态特征;(2)混合潜在融合:通过球面线性插值融合真实画作与合成代理图像的特征,确保过渡平滑并增强时间一致性。实验表明,该方法显著提升文本语义对齐度,同时忠实保留原画独特特征与完整性。关键优势在于无需任何模型训练或可学习参数,可即插即用集成至现有 I2V 方法中,是实现真实画作动画化的理想方案。更多动效示例见项目官网。

原文摘要 · Abstract (English)

We introduce a training-free framework specifically designed to bring real-world static paintings to life through image-to-video (I2V) synthesis, addressing the persistent challenge of aligning these motions with textual guidance while preserving fidelity to the original artworks. Existing I2V methods, primarily trained on natural video datasets, often struggle to generate dynamic outputs from static paintings. It remains challenging to generate motion while maintaining visual consistency with real-world paintings. This results in two distinct failure modes: either static outputs due to limited text-based motion interpretation or distorted dynamics caused by inadequate alignment with real-world artistic styles. We leverage the advanced text-image alignment capabilities of pre-trained image models to guide the animation process. Our approach introduces synthetic proxy images through two key innovations: (1) Dual-path score distillation: We employ a dual-path architecture to distill motion priors from both real and synthetic data, preserving static details from the original painting while learning dynamic characteristics from synthetic frames. (2) Hybrid latent fusion: We integrate hybrid features extracted from real paintings and synthetic proxy images via spherical linear interpolation in the latent space, ensuring smooth transitions and enhancing temporal consistency. Experimental evaluations confirm that our approach significantly improves semantic alignment with text prompts while faithfully preserving the unique characteristics and integrity of the original paintings. Crucially, by achieving enhanced dynamic effects without requiring any model training or learnable parameters, our framework enables plug-and-play integration with existing I2V methods, making it an ideal solution for animating real-world paintings. More animated examples can be found on our project website.

图像生成动画化无训练艺术修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。