arXiv:2508.03955cs.CV2025-08被引 2

用大量噪声视频训练音频对齐动画,大幅减少人工标注需求。

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm

  • 两阶段训练:先自动筛选海量噪声视频预训练,再小规模微调高质量数据。
  • 仅需1.9%额外参数即可实现音频条件控制,且不损失生成模型原有能力。
  • 新基准AVSync48含48类视频,多样性超以往3倍,适合开放世界应用。

近期音频同步视觉动画技术可通过特定类别的音频控制视频内容,但现有方法严重依赖昂贵的人工标注高质量、类别专属视频,在开放世界中难以扩展。本文提出一种高效的两阶段训练范式,利用丰富但嘈杂的视频实现音频同步动画的规模化。第一阶段自动筛选大规模视频用于预训练,使模型学习多样但不完美的音画对齐;第二阶段仅在少量人工标注的高质量样本上微调,显著降低人工成本。通过多特征条件输入与窗口注意力机制,增强每帧对音频上下文的感知。为高效训练,引入预训练文本到视频生成器和音频编码器,仅增加1.9%可训练参数即可学习音频条件能力,且不破坏生成器原有知识。评估方面,我们构建了包含48个类别的新基准AVSync48,其多样性是此前基准的3倍。大量实验表明,本方法将对人工标注的依赖降低超过10倍,同时能泛化至众多开放类别。

原文摘要 · Abstract (English)

Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manual curation of high-quality, class-specific training videos, posing challenges to scaling up to diverse audio-video classes in the open world. In this work, we propose an efficient two-stage training paradigm to scale up audio-synchronized visual animation using abundant but noisy videos. In stage one, we automatically curate large-scale videos for pretraining, allowing the model to learn diverse but imperfect audio-video alignments. In stage two, we finetune the model on manually curated high-quality examples, but only at a small scale, significantly reducing the required human effort. We further enhance synchronization by allowing each frame to access rich audio context via multi-feature conditioning and window attention. To efficiently train the model, we leverage pretrained text-to-video generator and audio encoders, introducing only 1.9\% additional trainable parameters to learn audio-conditioning capability without compromising the generator's prior knowledge. For evaluation, we introduce AVSync48, a benchmark with videos from 48 classes, which is 3$\times$ more diverse than previous benchmarks. Extensive experiments show that our method significantly reduces reliance on manual curation by over 10$\times$, while generalizing to many open classes.

音频对齐视频生成自动化训练少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。