用零训练成本生成人物与物体互动图像,人脸保真且身体协调。
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction Generation
- 结合通用扩散模型与个性化人脸模型,通过文本引导生成互动图
- 在多个数据集上实现高真实感,交互对齐度提升18.7%
- 适合需要快速生成个性化角色互动场景的创作者和设计师
我们提出PersonaHOI,一个无需训练和微调的框架,将通用StableDiffusion模型与个性化人脸扩散(PFD)模型融合,生成身份一致的人体-物体交互(HOI)图像。现有PFD模型虽进步显著,但常过度强调面部特征而牺牲全身一致性。PersonaHOI引入额外的StableDiffusion(SD)分支,由面向交互的文本输入引导,并在PFD分支中加入交叉注意力约束,在潜空间与残差层进行空间融合,从而在保留个性化面部细节的同时确保非面部区域的交互合理性。实验基于新提出的交互对齐度量进行验证,结果表明PersonaHOI在真实感和可扩展性方面均表现卓越,确立了个性化人脸互动生成的新标准。代码将开源于https://github.com/JoyHuYY1412/PersonaHOI。
原文摘要 · Abstract (English)
We introduce PersonaHOI, a training- and tuning-free framework that fuses a general StableDiffusion model with a personalized face diffusion (PFD) model to generate identity-consistent human-object interaction (HOI) images. While existing PFD models have advanced significantly, they often overemphasize facial features at the expense of full-body coherence, PersonaHOI introduces an additional StableDiffusion (SD) branch guided by HOI-oriented text inputs. By incorporating cross-attention constraints in the PFD branch and spatial merging at both latent and residual levels, PersonaHOI preserves personalized facial details while ensuring interactive non-facial regions. Experiments, validated by a novel interaction alignment metric, demonstrate the superior realism and scalability of PersonaHOI, establishing a new standard for practical personalized face with HOI generation. Our code will be available at https://github.com/JoyHuYY1412/PersonaHOI
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。