arXiv:2608.11752cs.CVcs.SD2026-08

统一框架实现说话视频音画身份同步替换,效果自然且生成高效。

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

论文配图:UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
图 1 · 摘自论文原文
  • 用统一扩散模型同时处理音画身份迁移,保持动作与时间同步。
  • 仅需3次去噪步即可完成生成,比传统方法提速10倍以上。
  • 适合需要高质量视频换脸换声的创作者和影视制作人员。

说话视频中的人物替换需协调外观与声音的转移,同时保留原始动作、场景、语言内容及音视频同步。现有方法分别优化音视频模态,难以保证一致性。本文提出UniSwap,首个支持流式联合音画身份替换的框架。给定源视频、参考图像和语音片段,UniSwap在单个音画扩散变换器中迁移参考外观与声调,同时保留源内容与动态。为解决跨身份对齐数据稀缺问题,提出交换-重建训练流程,从真实视频中移除音画身份并以原视频作为重建目标。基于双向主干网络,通过上下文预训练实现联合替换,条件流式适配实现块因果键值缓存生成,高效自强迫去噪机制将每块采样步数从30降至3。高效多LoRA切换使三个去噪角色共享同一冻结主干。特征RoPE分解确保缓存位置在训练范围内,支持稳定长视频推理。实验表明,该方法具备强音画同步性、优异的身份保留能力、高效的流式生成与稳定的长序列输出。

原文摘要 · Abstract (English)

Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.

音画替换扩散模型视频生成流式生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。