用智能代理系统自动修复文生图中的细微瑕疵,效果更准更可信。
Agentic Retoucher for Text-To-Image Generation
- 分三步模拟人类纠错:感知定位问题、推理判断原因、行动精准修补。
- 在27000个瑕疵标注上测试,修复准确率显著优于现有方法。
- 适合需要高质量图像生成的设计师和科研人员使用。
文生图扩散模型如SDXL和FLUX已实现惊人逼真度,但肢体、面部、文字等小尺度失真仍普遍存在。现有修复方法要么成本高需多次重生成,要么依赖视觉语言模型(VLM)且空间定位能力弱,易引发语义漂移和不可靠局部修改。为此,我们提出Agentic Retoucher,一种分层决策驱动的框架,将生成后修正重构为类人感知-推理-行动循环。具体包括:(1) 感知代理,利用文本-图像一致性线索学习上下文显著性,精确定位细粒度失真;(2) 推理代理,通过渐进式偏好对齐进行符合人类判断的诊断;(3) 行动代理,根据用户偏好自适应规划局部修复。该设计融合感知证据、语言推理与可控修正,形成统一自纠错流程。为支持细粒度监督与量化评估,我们构建了GenBlemish-27K数据集,包含6000张文生图图像及27000个跨12类的瑕疵标注区域。大量实验表明,Agentic Retoucher在感知质量、失真定位与人类偏好对齐方面均持续优于当前最优方法,确立了自修正、感知可靠的文生图生成新范式。
原文摘要 · Abstract (English)
Text-to-image (T2I) diffusion models such as SDXL and FLUX have achieved impressive photorealism, yet small-scale distortions remain pervasive in limbs, face, text and so on. Existing refinement approaches either perform costly iterative re-generation or rely on vision-language models (VLMs) with weak spatial grounding, leading to semantic drift and unreliable local edits. To close this gap, we propose Agentic Retoucher, a hierarchical decision-driven framework that reformulates post-generation correction as a human-like perception-reasoning-action loop. Specifically, we design (1) a perception agent that learns contextual saliency for fine-grained distortion localization under text-image consistency cues, (2) a reasoning agent that performs human-aligned inferential diagnosis via progressive preference alignment, and (3) an action agent that adaptively plans localized inpainting guided by user preference. This design integrates perceptual evidence, linguistic reasoning, and controllable correction into a unified, self-corrective decision process. To enable fine-grained supervision and quantitative evaluation, we further construct GenBlemish-27K, a dataset of 6K T2I images with 27K annotated artifact regions across 12 categories. Extensive experiments demonstrate that Agentic Retoucher consistently outperforms state-of-the-art methods in perceptual quality, distortion localization and human preference alignment, establishing a new paradigm for self-corrective and perceptually reliable T2I generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。