arXiv:2505.23742cs.CVcs.AI2025-05被引 16

解决任意参考视频生成中的身份混乱与主体混淆问题

MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement

  • 用区域感知掩码和通道拼接保留多主体外观特征
  • 引入语义注入的主体解耦机制,避免多参考主体混淆
  • 四阶段数据流程有效减少复制粘贴伪影,适合可控视频生成研究者

我们针对任意参考视频生成任务,旨在根据任意类型和组合的参考主体及文本提示生成视频。该任务面临身份不一致、多参考主体间纠缠以及复制粘贴伪影等持续挑战。为此,我们提出MAGREF,一个统一且高效的任意参考视频生成框架。方法结合掩码引导与主体解耦机制,实现对多样化参考图像和文本提示的灵活合成。具体而言,掩码引导采用区域感知掩码与像素级通道拼接,沿通道维度保留多主体外观特征,既保持身份一致性,又不改变预训练主干结构。为缓解主体混淆,引入主体解耦机制,将文本条件推导出的每个主体语义值注入其对应视觉区域。此外,构建四阶段数据流水线以生成多样训练对,有效减轻复制粘贴伪影。在综合基准上的大量实验表明,MAGREF持续优于现有最先进方法,为可扩展、可控且高保真的任意参考视频合成开辟道路。代码与模型见:https://github.com/MAGREF-Video/MAGREF。

原文摘要 · Abstract (English)

We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual prompts. This task faces persistent challenges, including identity inconsistency, entanglement among multiple reference subjects, and copy-paste artifacts. To address these issues, we introduce MAGREF, a unified and effective framework for any-reference video generation. Our approach incorporates masked guidance and a subject disentanglement mechanism, enabling flexible synthesis conditioned on diverse reference images and textual prompts. Specifically, masked guidance employs a region-aware masking mechanism combined with pixel-wise channel concatenation to preserve appearance features of multiple subjects along the channel dimension. This design preserves identity consistency and maintains the capabilities of the pre-trained backbone, without requiring any architectural changes. To mitigate subject confusion, we introduce a subject disentanglement mechanism which injects the semantic values of each subject derived from the text condition into its corresponding visual region. Additionally, we establish a four-stage data pipeline to construct diverse training pairs, effectively alleviating copy-paste artifacts. Extensive experiments on a comprehensive benchmark demonstrate that MAGREF consistently outperforms existing state-of-the-art approaches, paving the way for scalable, controllable, and high-fidelity any-reference video synthesis. Code and model can be found at: https://github.com/MAGREF-Video/MAGREF

视频生成扩散模型主体解耦可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。