arXiv:2605.23178cs.CV2026-05International Conf…被引 1

让多人互动场景生成更自然,通过姿态与图像协同迭代生成。

Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes

论文配图:Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes
图 1 · 摘自论文原文
  • 用姿态图和图像联合建模,让人体结构与外观一起演化。
  • 在多人群体场景中显著提升生成多样性与提示对齐度。
  • 适合需要复杂多人交互生成的视觉创作与角色设计。

尽管近期取得进展,文本到图像模型在生成语义多样且构图准确的多人互动场景时仍存在困难,常出现布局重复、刻板姿势和交互不自然等问题。本文提出一种双模态姿态-图像表示,将以人物为中心的结构先验引入预训练扩散变换器。模型联合预测2D姿态可视化图及其对应的RGB图像,使结构与外观在学习过程中协同演化。核心采用跨模态对齐机制,绑定文本、姿态与图像表征,确保多模态间的一致性。此外,设计迭代式场景构建方案,逐步生成复杂多人互动,有效分解整体生成复杂度。大量实验表明,该方法在多人图像生成中显著提升了提示对齐度与场景多样性。

原文摘要 · Abstract (English)

Despite recent progress, text-to-image models still struggle to generate semantically diverse and compositionally accurate multi-person interaction scenes, often collapsing to repetitive layouts, stereotypical poses, and poorly grounded interactions. In this work, we bridge this gap by introducing a dual pose-image representation that brings person-centric structural priors into pretrained diffusion transformers. Our model jointly predicts a 2D pose visualization image and its corresponding RGB image, enabling structure and appearance to co-evolve during learning. At its core, a cross-modal alignment scheme binds text, pose, and image representations, ensuring consistent grounding across modalities. Furthermore, we design an iterative scene construction scheme, progressively generating complex multi-human interactions while effectively decomposing the overall generation complexity. Extensive experiments demonstrate that our method substantially improves prompt alignment and scene diversity in multi-person image generation.

多人群体生成扩散模型姿态引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。