arXiv:2512.18614cs.CVcs.AI2025-12

纯文本生成高质量动画,效果超越现有方法。

PTTA: A Pure Text-to-Animation Framework for High-Quality Creation

  • 基于HunyuanVideo模型微调,构建纯文本到动画的生成框架。
  • 在自建高质量配对数据集上训练,生成动画更符合风格要求。
  • 适合需要高效生成动画内容的创作者和影视团队。

传统动画制作流程复杂且人工成本高。尽管Sora、Kling和CogVideoX等视频生成模型在自然视频合成上表现优异,但在动画生成方面仍存在明显局限。近期如AniSora通过微调图像到视频模型实现动画风格生成,但文本到视频方向的探索仍不足。本文提出PTTA,一个纯文本到动画的生成框架。我们首先构建了一个小规模但高质量的动画视频与文本描述配对数据集。在此基础上,基于预训练的文本到视频模型HunyuanVideo进行微调,使其适配动画风格生成。多维度视觉评估显示,该方法在动画视频合成上持续优于现有基线模型。

原文摘要 · Abstract (English)

Traditional animation production involves complex pipelines and significant manual labor cost. While recent video generation models such as Sora, Kling, and CogVideoX achieve impressive results on natural video synthesis, they exhibit notable limitations when applied to animation generation. Recent efforts, such as AniSora, demonstrate promising performance by fine-tuning image-to-video models for animation styles, yet analogous exploration in the text-to-video setting remains limited. In this work, we present PTTA, a pure text-to-animation framework for high-quality animation creation. We first construct a small-scale but high-quality paired dataset of animation videos and textual descriptions. Building upon the pretrained text-to-video model HunyuanVideo, we perform fine-tuning to adapt it to animation-style generation. Extensive visual evaluations across multiple dimensions show that the proposed approach consistently outperforms comparable baselines in animation video synthesis.

文本生成动画生成视频生成微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。