arXiv:2501.08505cs.CVeess.IV2025-01AAAI被引 25

Yuan自动修复生成图像缺陷,提升真实感和可用性。

Yuan: Yielding Unblemished Aesthetics Through A Unified Network for Visual Imperfections Removal in Generated Images

  • 结合文本提示与图像分割,自动生成需修复区域的精确掩码
  • 在ImageNet100和Stanford Dogs上实现更高NIQE、BRISQUE、PI得分
  • 无需人工标注,适合需要高质量图像的创意与科研场景

生成式AI在艺术创作与科学可视化等领域具有巨大潜力,但常因人体结构错误、物体位置不当及文字错位等视觉缺陷而降低实用性。为此,我们提出Yuan框架,可自动修正文本到图像生成中的视觉缺陷。Yuan同时利用文本提示与图像分割结果,自动生成需修复区域的精确掩码,无需人工干预——这是以往方法的常见限制。随后,先进的修复模块在识别出的区域中无缝融入语义连贯的内容,保持原图完整性与文本提示的一致性。在ImageNet100、Stanford Dogs及自建数据集上的大量实验表明,Yuan在定量指标(如NIQE、BRISQUE、PI)上均优于现有方法,并获得更优的定性评价。这些结果证明Yuan显著提升了生成图像的质量与实际应用价值。

原文摘要 · Abstract (English)

Generative AI presents transformative potential across various domains, from creative arts to scientific visualization. However, the utility of AI-generated imagery is often compromised by visual flaws, including anatomical inaccuracies, improper object placements, and misplaced textual elements. These imperfections pose significant challenges for practical applications. To overcome these limitations, we introduce \textit{Yuan}, a novel framework that autonomously corrects visual imperfections in text-to-image synthesis. \textit{Yuan} uniquely conditions on both the textual prompt and the segmented image, generating precise masks that identify areas in need of refinement without requiring manual intervention -- a common constraint in previous methodologies. Following the automated masking process, an advanced inpainting module seamlessly integrates contextually coherent content into the identified regions, preserving the integrity and fidelity of the original image and associated text prompts. Through extensive experimentation on publicly available datasets such as ImageNet100 and Stanford Dogs, along with a custom-generated dataset, \textit{Yuan} demonstrated superior performance in eliminating visual imperfections. Our approach consistently achieved higher scores in quantitative metrics, including NIQE, BRISQUE, and PI, alongside favorable qualitative evaluations. These results underscore \textit{Yuan}'s potential to significantly enhance the quality and applicability of AI-generated images across diverse fields.

图像修复生成模型视觉质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。