arXiv:2504.04903cs.CVcs.AI2025-04被引 18

统一多模态框架,100+低级视觉任务一键处理

Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision

  • 用文本与图像提示协同控制,支持任意分辨率生成
  • 在1K分辨率下表现最优,细节保真度高
  • 适合需要多任务通用图像处理的开发者

我们提出Lumina-OmniLV(简称OmniLV),一种面向低级视觉任务的通用多模态多任务框架,可解决图像修复、图像增强、弱语义密集预测和风格化四大类共100多个子任务。该框架利用文本和视觉提示实现灵活交互,基于扩散Transformer(DiT)生成先验,支持任意分辨率输入,在1K分辨率下达到最佳性能,同时保持精细细节和高保真度。大量实验表明,分别编码文本与视觉指令,并结合浅层特征控制进行联合训练,是缓解任务歧义、提升多任务泛化能力的关键。研究还发现,将高层生成任务融入低级视觉模型会损害细节敏感型恢复任务的表现。这些发现为构建更鲁棒、更具泛化能力的低级视觉系统提供了新思路。

原文摘要 · Abstract (English)

We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks across four major categories: image restoration, image enhancement, weak-semantic dense prediction, and stylization. OmniLV leverages both textual and visual prompts to offer flexible and user-friendly interactions. Built on Diffusion Transformer (DiT)-based generative priors, our framework supports arbitrary resolutions -- achieving optimal performance at 1K resolution -- while preserving fine-grained details and high fidelity. Through extensive experiments, we demonstrate that separately encoding text and visual instructions, combined with co-training using shallow feature control, is essential to mitigate task ambiguity and enhance multi-task generalization. Our findings also reveal that integrating high-level generative tasks into low-level vision models can compromise detail-sensitive restoration. These insights pave the way for more robust and generalizable low-level vision systems.

多模态图像修复扩散模型通用框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。