arXiv:2507.14801cs.CV2025-07被引 6

用视觉提示统一建模低级视觉任务,支持多任务泛化与可扩展训练。

Exploring Scalable Unified Modeling for General Low-Level Vision

  • 以输入输出图像对作为视觉提示,构建可灵活适配多种架构的统一框架。
  • 在超100个任务上训练,模型规模越大、任务越多,泛化能力越强,小数据任务受益显著。
  • 适用于零样本迁移、少样本微调等场景,适合需要多任务通用模型的研究者。

低级视觉涵盖图像修复、增强、风格化和特征提取等多种任务,其任务定义与输出域差异显著。为解决跨任务统一建模的挑战,我们提出视觉任务提示图像处理(VPIP)框架,利用输入-目标图像对作为视觉提示,引导模型完成多样化的低级视觉任务。该框架包含端到端图像处理主干、提示编码器与提示交互模块,支持灵活集成各类架构并有效利用任务特定视觉表示。基于此设计,我们构建统一的低级视觉模型GenLV,并在多个代表性任务上评估性能。为探索方法可扩展性,我们沿模型容量与任务多样性两个维度进行扩展,构建包含超过100个低级视觉任务的大规模基准,并训练多个不同规模的模型版本。实验表明,该方法在广泛任务上均表现优异。值得注意的是,训练任务数量增加能提升泛化能力,尤其对数据有限的任务效果明显,表明模型可通过联合训练学习可迁移表征。进一步在零样本泛化、少样本迁移与任务特定微调场景下的评估证实了模型的强大适应性,验证了该框架作为通用低级视觉建模基础的有效性、可扩展性与潜力。

原文摘要 · Abstract (English)

Low-level vision involves a wide spectrum of tasks, including image restoration, enhancement, stylization, and feature extraction, which differ significantly in both task formulation and output domains. To address the challenge of unified modeling across such diverse tasks, we propose a Visual task Prompt-based Image Processing (VPIP) framework that leverages input-target image pairs as visual prompts to guide the model in performing a variety of low-level vision tasks. The framework comprises an end-to-end image processing backbone, a prompt encoder, and a prompt interaction module, enabling flexible integration with various architectures and effective utilization of task-specific visual representations. Based on this design, we develop a unified low-level vision model, GenLV, and evaluate its performance across multiple representative tasks. To explore the scalability of this approach, we extend the framework along two dimensions: model capacity and task diversity. We construct a large-scale benchmark consisting of over 100 low-level vision tasks and train multiple versions of the model with varying scales. Experimental results show that the proposed method achieves considerable performance across a wide range of tasks. Notably, increasing the number of training tasks enhances generalization, particularly for tasks with limited data, indicating the model's ability to learn transferable representations through joint training. Further evaluations in zero-shot generalization, few-shot transfer, and task-specific fine-tuning scenarios demonstrate the model's strong adaptability, confirming the effectiveness, scalability, and potential of the proposed framework as a unified foundation for general low-level vision modeling.

低级视觉统一建模提示学习可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。