arXiv:2412.09614cs.CVcs.CL2024-12被引 5

用知识图谱提升罕见概念的图像生成与编辑能力

RAVEL: Rare Concept Generation and Editing via Graph-driven Relational Guidance

  • 通过知识图谱检索关系上下文,弥补文本生成中罕见概念缺失
  • 在三个新基准上超越现有方法,显著提升生成准确性和叙事一致性
  • 无需训练,兼容主流扩散模型,适合创意设计与文化表达场景

当前文本到图像扩散模型虽具备出色视觉质量,但在描绘罕见、复杂或文化内涵丰富的概念时仍受限于训练数据。我们提出RAVEL,一种无需训练的框架,通过将基于图的检索增强生成(RAG)集成到扩散流程中,显著提升罕见概念生成、上下文驱动图像编辑与自修正能力。不同于依赖视觉样例、静态描述或预训练知识的已有方法,RAVEL利用结构化知识图谱获取组合性、象征性与关系性上下文,实现无视觉先验情况下的精细语义锚定。为进一步优化生成质量,提出SRD自修正模块,通过多维度对齐反馈迭代更新提示词,增强属性准确性、叙事连贯性与语义保真度。该框架具备模型无关性,兼容Stable Diffusion XL、Flux与DALL-E 3等主流模型。我们在三个新构建的基准——MythoBench、Rare-Concept-1K与NovelBench上进行广泛评估,结果表明RAVEL在感知质量、对齐度及LLM作为裁判的指标上均持续优于现有最先进方法,为长尾领域中可控且可解释的T2I生成提供了可靠范式。

原文摘要 · Abstract (English)

Despite impressive visual fidelity, current text-to-image (T2I) diffusion models struggle to depict rare, complex, or culturally nuanced concepts due to training data limitations. We introduce RAVEL, a training-free framework that significantly improves rare concept generation, context-driven image editing, and self-correction by integrating graph-based retrieval-augmented generation (RAG) into diffusion pipelines. Unlike prior RAG and LLM-enhanced methods reliant on visual exemplars, static captions or pre-trained knowledge of models, RAVEL leverages structured knowledge graphs to retrieve compositional, symbolic, and relational context, enabling nuanced grounding even in the absence of visual priors. To further refine generation quality, we propose SRD, a novel self-correction module that iteratively updates prompts via multi-aspect alignment feedback, enhancing attribute accuracy, narrative coherence, and semantic fidelity. Our framework is model-agnostic and compatible with leading diffusion models including Stable Diffusion XL, Flux, and DALL-E 3. We conduct extensive evaluations across three newly proposed benchmarks - MythoBench, Rare-Concept-1K, and NovelBench. RAVEL also consistently outperforms SOTA methods across perceptual, alignment, and LLM-as-a-Judge metrics. These results position RAVEL as a robust paradigm for controllable and interpretable T2I generation in long-tail domains.

文本生成图像知识图谱罕见概念自修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。