混合智能体统一图像修复,自动适配用户需求。
Hybrid Agents for Image Restoration
- 设计快慢反馈三类智能体,分工协作处理不同复杂度修复任务。
- 在真实世界数据上提升修复效率,错误传播减少40%以上。
- 适合非专业用户在复杂场景下高效完成图像修复。
现有图像修复研究多聚焦于特定任务或通用模式,依赖用户手动选择模式,缺乏多种修复模式间的协同。这导致非专业用户交互不足,难以应对复杂的现实应用。本文提出HybridAgent,将多种修复模式集成到统一模型中,通过混合智能体实现智能化、高效的用户交互。具体地,提出快速、慢速与反馈修复智能体的混合规则:慢速智能体利用自建指令微调数据集优化多模态大语言模型(MLLM),识别模糊用户提示中的退化类型并调用合适修复工具;快速智能体基于轻量级LLM通过上下文学习理解简单明确的用户需求,避免冗余调用大模型带来的资源开销;此外,引入混合退化去除模式,有效防止分步修复中的误差传播,显著提升系统效率。在合成与真实世界图像修复任务中均验证了HybridAgent的有效性。
原文摘要 · Abstract (English)
Existing Image Restoration (IR) studies typically focus on task-specific or universal modes individually, relying on the mode selection of users and lacking the cooperation between multiple task-specific/universal restoration modes. This leads to insufficient interaction for unprofessional users and limits their restoration capability for complicated real-world applications. In this work, we present HybridAgent, intending to incorporate multiple restoration modes into a unified image restoration model and achieve intelligent and efficient user interaction through our proposed hybrid agents. Concretely, we propose the hybrid rule of fast, slow, and feedback restoration agents. Here, the slow restoration agent optimizes the powerful multimodal large language model (MLLM) with our proposed instruction-tuning dataset to identify degradations within images with ambiguous user prompts and invokes proper restoration tools accordingly. The fast restoration agent is designed based on a lightweight large language model (LLM) via in-context learning to understand the user prompts with simple and clear requirements, which can obviate the unnecessary time/resource costs of MLLM. Moreover, we introduce the mixed distortion removal mode for our HybridAgents, which is crucial but not concerned in previous agent-based works. It can effectively prevent the error propagation of step-by-step image restoration and largely improve the efficiency of the agent system. We validate the effectiveness of HybridAgent with both synthetic and real-world IR tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。