arXiv:2510.08508cs.CV2025-10被引 19

用三个智能代理自动完成视频修复,效果比现有方法更好。

MoA-VR: A Mixture-of-Agents System Towards All-in-One Video Restoration

  • 分三步:识别退化类型、选择修复方案、评估修复质量
  • 在多个数据集上指标优于现有方法,感知质量更优
  • 适合需要全自动修复复杂视频的工程师或研究人员

真实世界视频常因采集和传输条件多样而出现噪声、压缩伪影和低光等复杂退化。现有修复方法通常依赖专业人员手动选择专用模型,或使用单一架构,难以泛化。受专家经验启发,我们提出首个混合智能体视频修复系统 MoA-VR,通过三个协同代理模拟专业流程:退化识别、路由与修复、修复质量评估。我们构建了大规模高分辨率退化识别基准,并设计基于视觉语言模型(VLM)的退化识别器;引入由大语言模型(LLM)驱动的自适应路由模块,通过观察工具使用模式自主学习有效修复策略;为评估中间与最终结果质量,构建了恢复视频质量(Res-VQ)数据集,并设计专用于修复任务的VLM-based VQA模型。大量实验表明,MoA-VR能有效处理多种复合退化,在客观指标和主观感知质量上均显著优于现有基线,验证了多模态智能与模块化推理在通用视频修复中的潜力。

原文摘要 · Abstract (English)

Real-world videos often suffer from complex degradations, such as noise, compression artifacts, and low-light distortions, due to diverse acquisition and transmission conditions. Existing restoration methods typically require professional manual selection of specialized models or rely on monolithic architectures that fail to generalize across varying degradations. Inspired by expert experience, we propose MoA-VR, the first \underline{M}ixture-\underline{o}f-\underline{A}gents \underline{V}ideo \underline{R}estoration system that mimics the reasoning and processing procedures of human professionals through three coordinated agents: Degradation Identification, Routing and Restoration, and Restoration Quality Assessment. Specifically, we construct a large-scale and high-resolution video degradation recognition benchmark and build a vision-language model (VLM) driven degradation identifier. We further introduce a self-adaptive router powered by large language models (LLMs), which autonomously learns effective restoration strategies by observing tool usage patterns. To assess intermediate and final processed video quality, we construct the \underline{Res}tored \underline{V}ideo \underline{Q}uality (Res-VQ) dataset and design a dedicated VLM-based video quality assessment (VQA) model tailored for restoration tasks. Extensive experiments demonstrate that MoA-VR effectively handles diverse and compound degradations, consistently outperforming existing baselines in terms of both objective metrics and perceptual quality. These results highlight the potential of integrating multimodal intelligence and modular reasoning in general-purpose video restoration systems.

视频修复智能体系统视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。