arXiv:2603.02138cs.CV2026-03中稿 · CVPR

用多模态指令生成高质量矢量动画,支持灵活控制。

OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens

  • 将Lottie JSON转为结构化命令序列,便于模型学习
  • 基于预训练视觉语言模型,实现多模态指令精准生成
  • 构建200万条专业级矢量动画数据集,推动领域发展

OmniLottie是一种多功能框架,可从多模态指令生成高质量矢量动画。为实现灵活的运动与视觉内容控制,研究聚焦于轻量级的Lottie格式——一种用于表示形状和动画行为的JSON格式。然而,原始的Lottie JSON文件包含大量不变的结构元数据和格式标记,给矢量动画生成的学习带来挑战。为此,我们设计了一种高效的Lottie分词器,将JSON文件转换为表示形状、动画函数和控制参数的结构化命令与参数序列。该分词器使我们能够基于预训练视觉语言模型,遵循多模态交错指令生成高质量矢量动画。为进一步推动矢量动画生成研究,我们构建了MMLottie-2M数据集,包含200万条由专业人士设计的矢量动画及其文本与视觉标注。大量实验表明,OmniLottie能生成生动且语义对齐的矢量动画,严格遵循多模态人类指令。

原文摘要 · Abstract (English)

OmniLottie is a versatile framework that generates high quality vector animations from multi-modal instructions. For flexible motion and visual content control, we focus on Lottie, a light weight JSON formatting for both shapes and animation behaviors representation. However, the raw Lottie JSON files contain extensive invariant structural metadata and formatting tokens, posing significant challenges for learning vector animation generation. Therefore, we introduce a well designed Lottie tokenizer that transforms JSON files into structured sequences of commands and parameters representing shapes, animation functions and control parameters. Such tokenizer enables us to build OmniLottie upon pretrained vision language models to follow multi-modal interleaved instructions and generate high quality vector animations. To further advance research in vector animation generation, we curate MMLottie-2M, a large scale dataset of professionally designed vector animations paired with textual and visual annotations. With extensive experiments, we validate that OmniLottie can produce vivid and semantically aligned vector animations that adhere closely to multi modal human instructions.

矢量动画多模态生成Lottie生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。