让多人生成不混脸、姿势准,靠检索+解耦位置编码。
ASTRA: Enhancing Multi-Subject Generation with Retrieval-Augmented Pose Guidance and Disentangled Position Embedding

- 用检索获取干净姿态先验,解耦身份与姿态结构。
- 在复杂姿态基准上姿态遵循度达新高,身份保持清晰。
- 适合需要多角色精准动作生成的创意设计场景。
主体驱动图像生成在个性化内容创作中表现优异,但通常仅限于单主体和常见姿态。现有方法在处理多个主体且动作复杂的场景时,面临身份保留与姿态精确性之间的根本矛盾:外观与结构信号在模型中纠缠,导致身份融合与姿态扭曲。为此,我们提出ASTRA(Adaptive Synthesis through Targeted Retrieval Augmentation),一种在统一扩散Transformer架构中解耦主体外观与姿态结构的新框架。ASTRA采用双策略:首先通过检索增强姿态(RAG-Pose)管道从精选数据库获取清晰的结构先验;其次其核心生成模型使用增强型通用旋转位置编码(EURoPE),实现身份令牌与空间位置解耦,同时将姿态令牌绑定至画布。此外,解耦语义调制(DSM)适配器将身份保持任务移至文本条件流。大量实验表明,该集成方法实现了更优的解耦效果。在自建的基于COCO的复杂姿态基准上,ASTRA在姿态遵循度上达到新SOTA,同时在DreamBench上保持高身份保真度与文本对齐能力。
原文摘要 · Abstract (English)
Subject-driven image generation has shown great success in creating personalized content, but its capabilities are largely confined to single subjects in common poses. Current approaches face a fundamental conflict when handling multiple subjects with complex, distinct actions: preserving individual identities while enforcing precise pose structures. This challenge often leads to identity fusion and pose distortion, as appearance and structure signals become entangled within the model's architecture. To resolve this conflict, we introduce ASTRA(Adaptive Synthesis through Targeted Retrieval Augmentation), a novel framework that architecturally disentangles subject appearance from pose structure within a unified Diffusion Transformer. ASTRA achieves this through a dual-pronged strategy. It first employs a Retrieval-Augmented Pose (RAG-Pose) pipeline to provide a clean, explicit structural prior from a curated database. Then, its core generative model learns to process these dual visual conditions using our Enhanced Universal Rotary Position Embedding (EURoPE), an asymmetric encoding mechanism that decouples identity tokens from spatial locations while binding pose tokens to the canvas. Concurrently, a Disentangled Semantic Modulation (DSM) adapter offloads the identity preservation task into the text conditioning stream. Extensive experiments demonstrate that our integrated approach achieves superior disentanglement. On our designed COCO-based complex pose benchmark, ASTRA achieves a new state-of-the-art in pose adherence, while maintaining high identity fidelity and text alignment in DreamBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。