用可学习的关键点提升稀疏姿态信号的图像生成控制力
Rethink Sparse Signals for Pose-guided Text-to-image Generation
- 将OpenPose转为可学习的空间表征,增强关键点表达能力
- 通过关键点概念学习实现姿态对齐,生成更符合提示的图像
- 在动物和人体生成中表现优于同类稀疏方法,媲美密集信号
近期研究倾向使用密集信号(如深度图、DensePose)替代稀疏信号(如OpenPose)以提供更详细的姿态引导。然而密集表示带来编辑困难和与文本提示不一致等问题。本文重新审视稀疏信号的价值,因其结构简单且与形状无关,仍具潜力。提出新型Spatial-Pose ControlNet(SP-Ctrl),赋予稀疏信号更强的可控性。具体地,将OpenPose扩展为可学习的空间表示,使关键点嵌入更具判别性和表现力;同时引入关键点概念学习,促使关键点令牌关注各自空间位置,提升姿态对齐效果。在动物与人体图像生成任务上,SP-Ctrl在稀疏姿态引导下优于现有空间可控文本到图像生成方法,并达到密集信号方法的性能水平。此外,该方法在跨物种与多样化生成中展现良好能力。
原文摘要 · Abstract (English)
Recent works favored dense signals (e.g., depth, DensePose), as an alternative to sparse signals (e.g., OpenPose), to provide detailed spatial guidance for pose-guided text-to-image generation. However, dense representations raised new challenges, including editing difficulties and potential inconsistencies with textual prompts. This fact motivates us to revisit sparse signals for pose guidance, owing to their simplicity and shape-agnostic nature, which remains underexplored. This paper proposes a novel Spatial-Pose ControlNet(SP-Ctrl), equipping sparse signals with robust controllability for pose-guided image generation. Specifically, we extend OpenPose to a learnable spatial representation, making keypoint embeddings discriminative and expressive. Additionally, we introduce keypoint concept learning, which encourages keypoint tokens to attend to the spatial positions of each keypoint, thus improving pose alignment. Experiments on animal- and human-centric image generation tasks demonstrate that our method outperforms recent spatially controllable T2I generation approaches under sparse-pose guidance and even matches the performance of dense signal-based methods. Moreover, SP-Ctrl shows promising capabilities in diverse and cross-species generation through sparse signals. Codes will be available at https://github.com/DREAMXFAR/SP-Ctrl.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。