arXiv:2607.14595cs.CV2026-07

用极少量参数实现视频生成模型高效微调,训练更稳定。

MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation

论文配图:MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation
图 1 · 摘自论文原文
  • 用轻量软提示控制生成,参数减少一个数量级
  • 仅需不到1%可训练参数,性能媲美全量微调
  • 双空间奖励反馈机制,提升条件引导训练稳定性

大规模视频扩散模型虽生成效果优异,但全量微调代价高昂。现有参数高效微调方法在十亿级模型上存在两大缺陷:仍需大量可训练参数,且基于奖励的训练在条件引导任务中易受噪声干扰导致优化不稳。我们提出 MagicPrompt,一种超轻量框架,实现极致参数效率与稳定奖励优化。它首先采用注意力嵌入式提示微调,通过极少量软提示引导生成,同时保留预训练知识;进一步引入双空间奖励反馈优化,利用自监督隐空间目标改进条件引导的奖励训练。实验表明,MagicPrompt 仅需不足1%可训练参数,即达竞争性性能,并显著降低训练成本。

原文摘要 · Abstract (English)

Large-scale video diffusion models deliver strong generation performance, but full fine-tuning for downstream tasks incurs prohibitive computational costs. Existing parameter-efficient fine-tuning (PEFT) methods have two critical flaws on billion-scale models: they still require substantial trainable parameters, and reward-based training suffers from noise-induced optimization instability in condition-guided tasks. We propose MagicPrompt, a lightweight framework that achieves extreme parameter efficiency and stable reward optimization. It first adopts Attention-Embedded Prompt Tuning, which steers generation via lightweight soft prompts with orders of magnitude fewer parameters while preserving pre-trained knowledge. It further introduces Dual-Space Reward Feedback Optimization, which uses self-supervised latent objectives to improve condition-guided reward training. Experiments show MagicPrompt reaches competitive performance with less than 1% trainable parameters and notably reduces training costs.

视频生成提示微调高效训练扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。