让多人合影编辑更精准,保持人物外观一致且可灵活控制。
GroupDiff: Diffusion-based Group Portrait Editing

- 用数据引擎生成训练样本,覆盖增删改人等需求
- 通过骨骼和注意力注入保留人物外观,避免变形
- 用框定位控制编辑位置,支持灵活调整人物
多人合影编辑需求广泛,但因人物互动复杂、姿态多样而困难。本文提出GroupDiff,三项核心贡献:1)构建数据引擎,生成用于训练的成对编辑数据,覆盖多种编辑场景;2)引入人物图像与骨骼信息至注意力模块,实现外观一致性保持;3)利用人物边界框重加权注意力矩阵,实现跨人物特征精准注入。大量实验表明,GroupDiff在编辑可控性与原图保真度上均优于现有方法,支持高效、精确的多人合影修改。
原文摘要 · Abstract (English)
Group portrait editing is highly desirable since users constantly want to add a person, delete a person, or manipulate existing persons. It is also challenging due to the intricate dynamics of human interactions and the diverse gestures. In this work, we present GroupDiff, a pioneering effort to tackle group photo editing with three dedicated contributions: 1) Data Engine: Since there is no labeled data for group photo editing, we create a data engine to generate paired data for training. The training data engine covers the diverse needs of group portrait editing. 2) Appearance Preservation: To keep the appearance consistent after editing, we inject the images of persons from the group photo into the attention modules and employ skeletons to provide intra-person guidance. 3) Control Flexibility: Bounding boxes indicating the locations of each person are used to reweight the attention matrix so that the features of each person can be injected into the correct places. This inter-person guidance provides flexible manners for manipulation. Extensive experiments demonstrate that GroupDiff exhibits state-of-the-art performance compared to existing methods. GroupDiff offers controllability for editing and maintains the fidelity of the original photos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。