4KAgent能自动将任意图像无损提升至4K超清,支持多种场景。
4KAgent: Agentic Any Image to 4K Super-Resolution
- 采用智能代理架构,按输入图像特性定制修复方案。
- 在11类任务26个基准上均达新顶尖水平,含医学影像和AI生成图。
- 特别优化人脸细节,适合摄影、医疗及艺术创作用户使用。
我们提出4KAgent,一个统一的智能体超分辨率通用系统,可将任意图像无差别放大至4K(甚至更高,支持迭代)。该系统能将严重退化的低分辨率图像(如256x256)恢复为清晰逼真的4K输出。4KAgent包含三大核心组件:(1) 配置模块,根据具体应用定制处理流程;(2) 感知代理,结合视觉-语言模型与图像质量评估专家,分析输入并制定个性化修复计划;(3) 修复代理,遵循递归执行-反思机制,由质量驱动的专家混合策略指导每一步最优输出选择。此外,系统嵌入专用人脸修复管道,显著增强肖像与自拍中面部细节。我们在11类任务共26个多样化基准上进行严格评估,涵盖自然图像、人像、AI生成内容、卫星图像、荧光显微镜、眼底、超声、X光等医学影像,表现优于现有方法,在感知(如NIQE、MUSIQ)与保真度(如PSNR)指标上均领先。通过建立低层视觉任务的新一代智能体范式,推动视觉自主智能体在多领域的发展。代码、模型与结果将公开于:https://4kagent.github.io。
原文摘要 · Abstract (English)
We present 4KAgent, a unified agentic super-resolution generalist system designed to universally upscale any image to 4K resolution (and even higher, if applied iteratively). Our system can transform images from extremely low resolutions with severe degradations, for example, highly distorted inputs at 256x256, into crystal-clear, photorealistic 4K outputs. 4KAgent comprises three core components: (1) Profiling, a module that customizes the 4KAgent pipeline based on bespoke use cases; (2) A Perception Agent, which leverages vision-language models alongside image quality assessment experts to analyze the input image and make a tailored restoration plan; and (3) A Restoration Agent, which executes the plan, following a recursive execution-reflection paradigm, guided by a quality-driven mixture-of-expert policy to select the optimal output for each step. Additionally, 4KAgent embeds a specialized face restoration pipeline, significantly enhancing facial details in portrait and selfie photos. We rigorously evaluate our 4KAgent across 11 distinct task categories encompassing a total of 26 diverse benchmarks, setting new state-of-the-art on a broad spectrum of imaging domains. Our evaluations cover natural images, portrait photos, AI-generated content, satellite imagery, fluorescence microscopy, and medical imaging like fundoscopy, ultrasound, and X-ray, demonstrating superior performance in terms of both perceptual (e.g., NIQE, MUSIQ) and fidelity (e.g., PSNR) metrics. By establishing a novel agentic paradigm for low-level vision tasks, we aim to catalyze broader interest and innovation within vision-centric autonomous agents across diverse research communities. We will release all the code, models, and results at: https://4kagent.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。