TARO通过动态对齐与事件触发条件,提升视频转音频的音画同步与质量。
TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis
- 动态调整隐空间对齐强度,随噪声调度自适应优化
- 引入音符起始提示,精准捕捉视觉事件对应的音频起点
- 在两个数据集上实现97.19%同步准确率,音画误差降低超50%
本文提出Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning (TARO),一种用于高保真、时序一致视频转音频合成的新框架。基于流模型变压器(flow-based transformers),该框架具备稳定训练与连续变换能力,显著提升音画同步性与音频质量。TARO引入两项关键创新:(1) 时间步自适应表示对齐(TRA),根据噪声调度动态调整隐空间对齐强度,确保特征演化平滑并提升保真度;(2) 起始时刻感知条件生成(OAC),融合起始事件提示作为音频相关视觉时刻的锐利标记,增强与动态视觉事件的同步能力。在VGGSound和Landscape数据集上的大量实验表明,TARO优于现有方法,相对降低53%的弗雷歇距离(FD)、29%的弗雷歇音频距离(FAD),并达到97.19%的对齐准确率,彰显其在音频质量和同步精度上的卓越表现。
原文摘要 · Abstract (English)
This paper introduces Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning (TARO), a novel framework for high-fidelity and temporally coherent video-to-audio synthesis. Built upon flow-based transformers, which offer stable training and continuous transformations for enhanced synchronization and audio quality, TARO introduces two key innovations: (1) Timestep-Adaptive Representation Alignment (TRA), which dynamically aligns latent representations by adjusting alignment strength based on the noise schedule, ensuring smooth evolution and improved fidelity, and (2) Onset-Aware Conditioning (OAC), which integrates onset cues that serve as sharp event-driven markers of audio-relevant visual moments to enhance synchronization with dynamic visual events. Extensive experiments on the VGGSound and Landscape datasets demonstrate that TARO outperforms prior methods, achieving relatively 53% lower Frechet Distance (FD), 29% lower Frechet Audio Distance (FAD), and a 97.19% Alignment Accuracy, highlighting its superior audio quality and synchronization precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。