用社交上下文反馈优化个性化图像生成,提升姿态与身份一致性
Improving Personalized Image Generation through Social Context Feedback
- 引入姿态、交互、人脸、视线等检测器进行反馈微调
- 在三个基准数据集上显著提升人物姿态、身份与视线一致性
- 按时间步动态融合低层(姿态)与高层(视线)反馈信号
个性化图像生成通过参考图像生成符合场景描述的特定人物图像,但存在三大缺陷:复杂动作(如“男性推摩托车”)姿态不准确,参考人物身份丢失,生成视线方向自然且与场景不符。本文提出基于反馈的微调方法,利用先进的姿态、人机交互、人脸识别和视线估计检测器,对扩散模型进行优化。设计分时步的反馈模块注入机制,根据信号层级(低层姿态或高层视线)动态调整。实验表明,该方法在三个基准数据集上显著改善了人物交互、身份保留与图像质量。
原文摘要 · Abstract (English)
Personalized image generation, where reference images of one or more subjects are used to generate their image according to a scene description, has gathered significant interest in the community. However, such generated images suffer from three major limitations -- complex activities, such as $<$man, pushing, motorcycle$>$ are not generated properly with incorrect human poses, reference human identities are not preserved, and generated human gaze patterns are unnatural/inconsistent with the scene description. In this work, we propose to overcome these shortcomings through feedback-based fine-tuning of existing personalized generation methods, wherein, state-of-art detectors of pose, human-object-interaction, human facial recognition and human gaze-point estimation are used to refine the diffusion model. We also propose timestep-based inculcation of different feedback modules, depending upon whether the signal is low-level (such as human pose), or high-level (such as gaze point). The images generated in this manner show an improvement in the generated interactions, facial identities and image quality over three benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。