arXiv:2602.20951cs.CVcs.AI2026-02

用智能代理自动生成带瑕疵的图像,提升视觉模型对艺术缺陷的理解能力。

See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis

  • 通过三个智能代理自动合成真实与带瑕疵图像对,无需人工标注。
  • 生成10万张带丰富瑕疵标注的图像,支持多种下游应用测试。
  • 适合研究扩散模型缺陷、图像修复和自动化数据构建的研究者。

尽管扩散模型取得进展,生成图像仍常含影响真实感的视觉瑕疵。单纯加大模型或更充分预训练无法保证完全消除瑕疵,因此瑕疵缓解成为关键研究方向。以往方法依赖人工标注的瑕疵数据集,成本高且难以扩展,亟需自动化获取瑕疵标注数据的方案。本文提出ArtiAgent,高效生成真实图像与注入瑕疵的图像对。其包含三个代理:感知代理识别并定位真实图像中的实体与子实体;合成代理通过扩散变换器中的新型逐块嵌入操作,利用瑕疵注入工具引入瑕疵;校验代理筛选合成瑕疵,并为每例生成局部与全局解释。基于ArtiAgent,我们合成10万张带丰富瑕疵标注的图像,并在多种应用中验证了其有效性与通用性。代码已公开。

原文摘要 · Abstract (English)

Despite recent advances in diffusion models, AI generated images still often contain visual artifacts that compromise realism. Although more thorough pre-training and bigger models might reduce artifacts, there is no assurance that they can be completely eliminated, which makes artifact mitigation a highly crucial area of study. Previous artifact-aware methodologies depend on human-labeled artifact datasets, which are costly and difficult to scale, underscoring the need for an automated approach to reliably acquire artifact-annotated datasets. In this paper, we propose ArtiAgent, which efficiently creates pairs of real and artifact-injected images. It comprises three agents: a perception agent that recognizes and grounds entities and subentities from real images, a synthesis agent that introduces artifacts via artifact injection tools through novel patch-wise embedding manipulation within a diffusion transformer, and a curation agent that filters the synthesized artifacts and generates both local and global explanations for each instance. Using ArtiAgent, we synthesize 100K images with rich artifact annotations and demonstrate both efficacy and versatility across diverse applications. Code is available at link.

视觉瑕疵扩散模型数据合成智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。