用一张或多张参考图,给目标人像换表情或发型等属性并生成动画。
Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer
- 用同一视频中两帧当伪配对,自监督学习跨身份属性迁移。
- 在扩散模型中通过空间注意力融合双参考特征,实现精准属性转移。
- 支持多属性组合与平滑插值,适合可控人像动画生成。
我们提出Durian,首个基于一个或多个参考图像实现跨身份属性迁移的人像动画生成方法。传统方法需同一个人的成对属性数据,但此类数据难获取。为此,我们设计自重建训练框架,利用普通人像视频进行无配对数据学习:同一视频中的两帧构成伪配对,一帧作为属性参考,另一帧作为身份参考。为支持该训练,我们引入双参考网络(Dual ReferenceNet),分别处理两个参考,并通过空间注意力在扩散模型中融合特征。为确保每条参考专用于身份或属性信息,我们对参考图施加互补掩码。两项机制协同使模型能重建原视频,自然学习跨身份属性迁移。为弥合自重建训练与跨身份推理之间的差距,我们提出掩码扩展策略与增强方案,提升对不同空间范围和错位情况下的属性迁移鲁棒性。Durian在人像动画属性迁移任务上达到当前最优性能,其双参考设计还支持单次生成内完成多属性组合与平滑插值,实现高度灵活可控的合成。
原文摘要 · Abstract (English)
We present Durian, the first method for generating portrait animation videos with cross-identity attribute transfer from one or more reference images to a target portrait. Training such models typically requires attribute pairs of the same individual, which are rarely available at scale. To address this challenge, we propose a self-reconstruction formulation that leverages ordinary portrait videos to learn attribute transfer without explicit paired data. Two frames from the same video act as a pseudo pair: one serves as an attribute reference and the other as an identity reference. To enable this self-reconstruction training, we introduce a Dual ReferenceNet that processes the two references separately and then fuses their features via spatial attention within a diffusion model. To make sure each reference functions as a specialized stream for either identity or attribute information, we apply complementary masking to the reference images. Together, these two components guide the model to reconstruct the original video, naturally learning cross-identity attribute transfer. To bridge the gap between self-reconstruction training and cross-identity inference, we introduce a mask expansion strategy and augmentation schemes, enabling robust transfer of attributes with varying spatial extent and misalignment. Durian achieves state-of-the-art performance on portrait animation with attribute transfer. Moreover, its dual reference design uniquely supports multi-attribute composition and smooth attribute interpolation within a single generation pass, enabling highly flexible and controllable synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。