arXiv:2509.21893cs.CV2025-09被引 5

用音频精准控制视频动作,生成更同步的高分辨率动态画面

Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers

  • 基于扩散模型,引入运动感知损失和音频同步引导机制
  • 在AVSync15和The Greatest Hits数据集上同步精度显著提升
  • 适合需要精确时序控制的音乐视频、动画生成场景

文本到视频和图像到视频生成在视觉质量上进展迅速,但在运动时机控制上仍受限。音频与视频动作在时间上对齐,是实现时序可控视频生成的有力条件。然而,现有音频到视频(A2V)模型因间接条件机制或有限的时间建模能力,难以实现细粒度同步。本文提出Syncphony,可生成380x640分辨率、24fps的视频,并与多样音频输入保持同步。方法基于预训练视频骨干网络,引入两个关键组件:(1) 运动感知损失,强化高运动区域的学习;(2) 音频同步引导,通过一个无音频层但视觉对齐的异步模型,在推理时更好利用音频线索,同时保持视觉质量。为评估同步性,提出CycleSync——一种基于视频到音频的度量,衡量生成视频中运动线索重建原始音频的能力。在AVSync15和The Greatest Hits数据集上的实验表明,Syncphony在同步精度和视觉质量上均优于现有方法。

原文摘要 · Abstract (English)

Text-to-video and image-to-video generation have made rapid progress in visual quality, but they remain limited in controlling the precise timing of motion. In contrast, audio provides temporal cues aligned with video motion, making it a promising condition for temporally controlled video generation. However, existing audio-to-video (A2V) models struggle with fine-grained synchronization due to indirect conditioning mechanisms or limited temporal modeling capacity. We present Syncphony, which generates 380x640 resolution, 24fps videos synchronized with diverse audio inputs. Our approach builds upon a pre-trained video backbone and incorporates two key components to improve synchronization: (1) Motion-aware Loss, which emphasizes learning at high-motion regions; (2) Audio Sync Guidance, which guides the full model using a visually aligned off-sync model without audio layers to better exploit audio cues at inference while maintaining visual quality. To evaluate synchronization, we propose CycleSync, a video-to-audio-based metric that measures the amount of motion cues in the generated video to reconstruct the original audio. Experiments on AVSync15 and The Greatest Hits datasets demonstrate that Syncphony outperforms existing methods in both synchronization accuracy and visual quality. Project page is available at: https://jibin86.github.io/syncphony_project_page

音频生成视频时序同步扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。