通过噪声感知纠错提升长视频生成的稳定性与音画同步性
Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification

- 根据噪声水平匹配残差注入,缓解自回归生成中的误差累积
- 在ST-Bench上实现更少的视觉漂移和更好的音画一致性
- 支持多镜头、多主体、参考图像引导的长时序音视频生成
自回归续写通过反复扩展短窗生成器,实现分钟级音视频生成。但模型训练使用真实历史数据,推理时依赖自身生成的历史,导致误差累积引发身份漂移、过度平滑及音画不同步。现有方法利用预测残差作为合成噪声来减少这种偏差,但我们发现残差修正效果高度依赖其产生的流匹配噪声水平。为此提出Vorch-Director,一种噪声水平感知的残差纠正策略:将每个残差与其原始噪声水平关联,并在训练中注入来自相同噪声区间的残差。通过使注入误差与去噪过程对齐,显著提升自回归历史的真实性,同时保持高效教师强制训练。基于音频-视觉LTX-2扩散变压器,Vorch-Director引入任务嵌入以区分历史视频、参考图像与目标视频,实现长时程生成的统一条件输入。结合干净条件接收器与混合任务训练,支持多镜头、多主体、参考引导的长视频生成。我们在ST-Bench上评估该方法,并引入一个新的长时程音视频基准,包含质量漂移与长程一致性指标。大量实验表明,相比强基线,其在稳定性与音画保真度上均有显著提升。
原文摘要 · Abstract (English)
Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated video and audio. However, models are trained on clean ground-truth histories, while inference relies on their own generated histories, where accumulated errors cause identity drift, over-smoothing, and audio-visual desynchronization. Recent methods reduce this mismatch by reusing prediction residuals as synthetic corruption, but we observe that the effectiveness of residual correction critically depends on the flow-matching noise level at which residuals are produced. We propose Vorch-Director, a noise-level-aware residual correction strategy that associates each residual with its originating noise level and injects residuals from matched noise regimes during training. By aligning injected errors with the denoising process, Vorch-Director produces more realistic autoregressive histories while retaining efficient teacher-forcing training. Built on the audio-visual LTX-2 diffusion transformer, Vorch-Director further introduces task embeddings to distinguish historical video, reference images, and target video, enabling unified conditioning for long-horizon generation. Together with a clean conditioning sink and mixed-task training, Vorch-Director supports multi-shot, multi-subject, reference-guided audio-visual long-video generation. We evaluate Vorch-Director on ST-Bench and introduce a new long-horizon audio-visual benchmark with metrics for quality drift and long-range consistency. Extensive experiments demonstrate improved stability and audio-visual fidelity over strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。