arXiv:2503.17728cs.CVcs.AI2025-03AAAI被引 4

单图个性化生成动态交互人物,保持身份一致

DynASyn: Multi-Subject Personalization Enabling Dynamic Action Synthesis

  • 用概念先验对齐主体外观与动作,防止过拟合
  • 仅需一张参考图即可生成多样动态行为图像
  • 适合需要个性化的视频生成与虚拟角色设计

近期文本到图像扩散模型的发展推动了个性化图像生成研究,即在参考图像中定制化合成特定主体。尽管现有方法可改变主体位置或同时个性化多个主体,但在修改主体行为或其动态交互方面仍存在困难,主要源于对参考图像的过拟合,尤其当仅有一张参考图时更为严重。本文提出DynASyn,一种从单张参考图实现多主体个性化的有效方法,解决了上述挑战。DynASyn通过将基于概念的先验与主体外观及动作对齐,保持个性化过程中的主体身份一致性,具体通过正则化主体标记与图像间的注意力图实现。此外,我们提出基于概念的提示与图像增强策略,以在身份保持与动作多样性之间取得更好平衡。采用基于SDE的编辑方法,结合增强后的提示生成多样化外观与动作,同时确保增强图像中身份的一致性。实验表明,DynASyn能够生成具有新场景和动态环境互动的高真实感图像,在定量与定性评估上均优于基线方法。

原文摘要 · Abstract (English)

Recent advances in text-to-image diffusion models spurred research on personalization, i.e., a customized image synthesis, of subjects within reference images. Although existing personalization methods are able to alter the subjects' positions or to personalize multiple subjects simultaneously, they often struggle to modify the behaviors of subjects or their dynamic interactions. The difficulty is attributable to overfitting to reference images, which worsens if only a single reference image is available. We propose DynASyn, an effective multi-subject personalization from a single reference image addressing these challenges. DynASyn preserves the subject identity in the personalization process by aligning concept-based priors with subject appearances and actions. This is achieved by regularizing the attention maps between the subject token and images through concept-based priors. In addition, we propose concept-based prompt-and-image augmentation for an enhanced trade-off between identity preservation and action diversity. We adopt an SDE-based editing guided by augmented prompts to generate diverse appearances and actions while maintaining identity consistency in the augmented images. Experiments show that DynASyn is capable of synthesizing highly realistic images of subjects with novel contexts and dynamic interactions with the surroundings, and outperforms baseline methods in both quantitative and qualitative aspects.

个性化生成动态交互单图建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。