arXiv:2503.17735cs.MMcs.CV2025-03被引 2

针对低资源动画贴纸生成,提出高效训练框架RDTF

RDTF: Resource-efficient Dual-mask Training Framework for Multi-frame Animated Sticker Generation

  • 用双掩码策略提升有限数据利用率与多样性
  • 在百万级数据集上训练小模型,效果优于主流微调方法
  • 适合资源受限场景下的动画贴纸生成任务

近年来视频生成技术取得显著进展,但将其应用于低帧率、抽象语义和长尾帧长分布的多帧动画贴纸生成(ASG)等资源受限任务仍具挑战。现有参数高效微调(PEFT)方法(如Adapter、LoRA)存在拟合能力不足和源域知识干扰问题。本文提出面向资源受限场景的高效双掩码训练框架(RDTF),通过三个核心设计实现从零开始训练紧凑模型:1)专为低帧率ASG优化的离散帧生成网络(DFGN),保障参数效率;2)基于双掩码的数据利用策略,增强有限数据的可用性与多样性;3)难度自适应课程学习方法,将样本熵分解为静态与动态成分,实现由易到难的训练收敛。为支持RDTF的端到端训练,我们构建了包含百万级样本的多模态动画贴纸数据集VSD2M,涵盖静态/动态贴纸及以动作为核心的文本描述,填补了ASG领域专用数据空白。实验表明,RDTF在定量与定性指标上均优于当前最优的PEFT方法(如I2V-Adapter、SimDA),验证了其在资源约束下的可行性与优越性。

原文摘要 · Abstract (English)

Recently, significant advancements have been achieved in video generation technology, but applying it to resource-constrained downstream tasks like multi-frame animated sticker generation (ASG) characterized by low frame rates, abstract semantics, and long tail frame length distribution-remains challenging. Parameter-efficient fine-tuning (PEFT) techniques (e.g., Adapter, LoRA) for large pre-trained models suffer from insufficient fitting ability and source-domain knowledge interference. In this paper, we propose Resource-Efficient Dual-Mask Training Framework (RDTF), a dedicated solution for multi-frame ASG task under resource constraints. We argue that training a compact model from scratch with million-level samples outperforms PEFT on large models, with RDTF realizing this via three core designs: 1) a Discrete Frame Generation Network (DFGN) optimized for low-frame-rate ASG, ensuring parameter efficiency; 2) a dual-mask based data utilization strategy to enhance the availability and diversity of limited data; 3) a difficulty-adaptive curriculum learning method that decomposes sample entropy into static and adaptive components, enabling easy-to-difficult training convergence. To provide high-quality data support for RDTFs training from scratch, we construct VSD2M-a million-level multi-modal animated sticker dataset with rich annotations (static and animated stickers, action-focused text descriptions)-filling the gap of dedicated animated data for ASG task. Experiments demonstrate that RDTF is quantitatively and qualitatively superior to state-of-the-art PEFT methods (e.g., I2V-Adapter, SimDA) on ASG tasks, verifying the feasibility of our framework under resource constraints.

动画贴纸高效训练数据增强小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。