arXiv:2511.21136cs.CVcs.AI2025-11被引 2

通过熵引导的渐进训练,显著降低人体视频生成的显存和时间开销。

Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning

  • 用条件熵膨胀评估模块重要性,优先训练关键组件。
  • 自适应提升计算复杂度,实现最高2.2倍加速与2.4倍显存减少。
  • 适合追求高效训练的人体视频生成研究者使用。

人体视频生成得益于扩散模型的发展迅速进步,但高分辨率多帧数据训练带来的高昂计算成本与巨大显存消耗仍是重大挑战。本文提出熵引导的优先渐进学习(Ent-Prog)框架,专为扩散模型在人体视频生成中的高效训练设计。首先,引入条件熵膨胀(CEI)评估不同模型组件对目标条件生成任务的重要性,实现关键组件的优先训练;其次,设计自适应渐进调度机制,通过衡量收敛效率动态增加训练复杂度。大量实验在三个数据集上验证了该方法的有效性,结果表明,Ent-Prog在保持生成性能的前提下,最多实现2.2倍训练速度提升和2.4倍GPU显存减少。

原文摘要 · Abstract (English)

Human video generation has advanced rapidly with the development of diffusion models, but the high computational cost and substantial memory consumption associated with training these models on high-resolution, multi-frame data pose significant challenges. In this paper, we propose Entropy-Guided Prioritized Progressive Learning (Ent-Prog), an efficient training framework tailored for diffusion models on human video generation. First, we introduce Conditional Entropy Inflation (CEI) to assess the importance of different model components on the target conditional generation task, enabling prioritized training of the most critical components. Second, we introduce an adaptive progressive schedule that adaptively increases computational complexity during training by measuring the convergence efficiency. Ent-Prog reduces both training time and GPU memory consumption while maintaining model performance. Extensive experiments across three datasets, demonstrate the effectiveness of Ent-Prog, achieving up to 2.2$\times$ training speedup and 2.4$\times$ GPU memory reduction without compromising generative performance.

视频生成扩散模型高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。