arXiv:2502.11158cs.CV2025-02被引 3

用左提示统一解决多种视觉任务,只需少量数据就能高效生成。

AnyRefill: A Unified, Data-Efficient Framework for Left-Prompt-Guided Vision Tasks

  • 用左右拼接构造上下文输入,让文本到图像模型适应多类视觉任务。
  • 仅需极少微调即可在生成、感知和编辑任务上达到顶尖性能。
  • 适合需要快速适配新任务的研究者或开发者,尤其擅长低资源场景。

本文提出一种新的左提示引导(LPG)范式,以应对多样化的基于参考的视觉任务。受人类创作过程启发,我们采用左右拼接方式构建上下文输入。在此基础上,提出 AnyRefill,作为 LeftRefill 的扩展,有效将基于 Diffusion Transformer(DiT)架构的文本到图像(T2I)模型适配至多种视觉任务。AnyRefill 利用先进 T2I 模型的修复先验,并引入灵活组件增强能力。通过结合任务特定的 LoRA 与拼接输入,无需额外视觉编码器即可在条件生成、视觉感知和图像编辑等任务中发挥潜力。该方法展现出显著的数据效率,仅需极小规模的任务微调即可保持高生成性能。大量消融实验表明,AnyRefill 超越其他图像条件注入方法,在多个任务上达到与开源前沿方法相当的结果。尤为突出的是,其表现可媲美 IC-Light、SeedEdit 等先进商业工具,即使在复杂场景下亦然。跨多样化任务的全面实验与分析验证了 LPG 公式简洁而强大的生成能力,确立 AnyRefill 为统一且高度数据高效的参考引导视觉任务解决方案。

原文摘要 · Abstract (English)

In this paper, we present a novel Left-Prompt-Guided (LPG) paradigm to address a diverse range of reference-based vision tasks. Inspired by the human creative process, we reformulate these tasks using a left-right stitching formulation to construct contextual input. Building upon this foundation, we propose AnyRefill, an extension of LeftRefill, that effectively adapts Text-to-Image (T2I) models to various vision tasks. AnyRefill leverages the inpainting priors of advanced T2I model based on the Diffusion Transformer (DiT) architecture, and incorporates flexible components to enhance its capabilities. By combining task-specific LoRAs with the stitching input, AnyRefill unlocks its potential across diverse tasks, including conditional generation, visual perception, and image editing, without requiring additional visual encoders. Meanwhile, AnyRefill exhibits remarkable data efficiency, requiring minimal task-specific fine-tuning while maintaining high generative performance. Through extensive ablation studies, we demonstrate that AnyRefill outperforms other image condition injection methods and achieves competitive results compared to state-of-the-art open-source methods. Notably, AnyRefill delivers results comparable to advanced commercial tools, such as IC-Light and SeedEdit, even in challenging scenarios. Comprehensive experiments and ablation studies across versatile tasks validate the strong generation of the proposed simple yet effective LPG formulation, establishing AnyRefill as a unified, highly data-efficient solution for reference-based vision tasks.

视觉任务文本生成数据效率扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。