解决服装生成中细节失真问题,提升纹理真实感。
RAGDiffusion++: From Macro-Retrieval to Micro-Fidelity Alignment for Garment Generation

- 通过检索增强与对抗正则化,对齐宏观结构与微观细节
- 新奖励模型在50万图像上训练,人类偏好准确率达84.67%
- 适合追求高精度服装生成的工业级应用开发者
标准服装资产生成——从多样的现实场景中恢复正面对应的平铺服装图像——具有巨大商业价值,但需同时保证宏观拓扑准确性与微观物理真实性。尽管先前工作RAGDiffusion通过检索增强的宏观约束有效消除大规模结构幻觉,但实现工业级微观纹理真实感仍是未解难题。我们首次将此瓶颈定义为高频轨迹坍缩:监督微调收敛于训练分布的条件均值,而该分布以平滑低频纹理为主,导致高频图案(如布料编织、复杂徽标)几乎无法采样。直接在训练后使用强化学习进一步引发伪影劫持,模型利用通用奖励模型中的语义偏差生成欺骗性棋盘噪声。核心洞见是:强化学习可从根本上重塑流模型的采样分布,在准确奖励引导下提升高保真轨迹概率,同时对抗正则化防止对奖励盲点的滥用。实现该原则需三要素:(i)内在能力,基于27,725对高复杂度服装数据集STGarment-Plus与双图像流FLUX架构升级;(ii)感知奖励,由在50万张图像上通过细粒度对比学习训练的新属性感知奖励模型Garment-RM提供,人类偏好准确率达84.67%;(iii)劫持防范,通过集成动态判别器到强化学习采样轨迹的对抗正则化GRPO(AR-GRPO)策略实现,既惩罚伪影又丰富真实高频细节。
原文摘要 · Abstract (English)
Standard clothing asset generation---restoring forward-facing flat-lay garment images from diverse real-world contexts---holds immense commercial value yet demands both macroscopic topological accuracy and microscopic physical fidelity. Although our previous work RAGDiffusion effectively eradicated large-scale structural hallucinations via retrieval-augmented macro-constraints, achieving industrial-grade micro-texture realism remains an unsolved bottleneck. We formally identify this limitation as High-Frequency Trajectory Collapse: supervised fine-tuning (SFT) converges to the conditional mean of the training distribution, which is dominated by smooth, low-frequency textures, causing high-frequency patterns (e.g., fabric weaves, intricate logos) to become nearly un-sampleable. Naively applying Reinforcement Learning (RL) post-training further triggers Artifact Hacking, where models exploit semantic biases in generic reward models by generating deceptive checkerboard noise. Our key insight is that RL can fundamentally reshape the sampling distribution of flow models---elevating the probability of high-fidelity trajectories under accurate reward guidance---while adversarial regularization prevents exploitation of reward blind spots. Realizing this principle requires three prerequisites: (i)inherent capacity, established through a 27,725-pair high-complexity garment dataset (STGarment-Plus) and a Dual-Image-Stream FLUX architecture upgrade; (ii)perceptive reward, provided by a novel attribute-aware reward model (Garment-RM) trained on 500K images via fine-grained contrastive learning, achieving 84.67% human preference accuracy; and (iii)hacking prevention, enforced by our Adversarial-Regularized GRPO (AR-GRPO) strategy that integrates a dynamic discriminator into the RL sampling trajectory to penalize artifacts while enriching authentic high-frequency details.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。