用外部知识增强生成,让衣服图像更真实、结构更准确。
RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation
- 引入检索增强机制,融合语言模型与数据库知识
- 在复杂场景下显著减少结构幻觉和纹理失真
- 适合需要高保真服装生成的工业级应用
标准服装资产生成需从多样化现实场景中恢复正向平铺的服装图像,背景清晰,但面临高度标准化的结构采样分布与复杂场景中语义缺失的挑战。现有模型空间感知能力有限,常在高规格生成任务中出现结构幻觉与纹理扭曲。为此,本文提出新型检索增强生成框架RAGDiffusion,通过整合语言模型与外部数据库知识,提升结构确定性并缓解幻觉。该框架包含两个阶段:(1) 基于对比学习与结构局部线性嵌入(SLLE)的结构聚合,提取全局结构与空间关键点,提供软硬双重引导以应对结构歧义;(2) 全层级保真服装生成,采用粗到细的纹理对齐策略,在扩散过程中保障图案与细节的一致性。在多个挑战性真实数据集上的实验表明,RAGDiffusion能生成结构与纹理高度忠实的服装资产,性能显著优于基线,是首个利用RAG实现高规格保真生成、对抗内在幻觉的开创性工作。
原文摘要 · Abstract (English)
Standard clothing asset generation involves restoring forward-facing flat-lay garment images displayed on a clear background by extracting clothing information from diverse real-world contexts, which presents significant challenges due to highly standardized structure sampling distributions and clothing semantic absence in complex scenarios. Existing models have limited spatial perception, often exhibiting structural hallucinations and texture distortion in this high-specification generative task. To address this issue, we propose a novel Retrieval-Augmented Generation (RAG) framework, termed RAGDiffusion, to enhance structure determinacy and mitigate hallucinations by assimilating knowledge from language models and external databases. RAGDiffusion consists of two processes: (1) Retrieval-based structure aggregation, which employs contrastive learning and a Structure Locally Linear Embedding (SLLE) to derive global structure and spatial landmarks, providing both soft and hard guidance to counteract structural ambiguities; and (2) Omni-level faithful garment generation, which introduces a coarse-to-fine texture alignment that ensures fidelity in pattern and detail components within the diffusing. Extensive experiments on challenging real-world datasets demonstrate that RAGDiffusion synthesizes structurally and texture-faithful clothing assets with significant performance improvements, representing a pioneering effort in high-specification faithful generation with RAG to confront intrinsic hallucinations and enhance fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。