arXiv:2603.22998cs.CV2026-03

智能视频修复代理,能精准识别退化类型并快速找到最优修复路径。

VQ-Jarvis: Retrieval-Augmented Video Restoration Agent with Sharp Vision and Fast Thought

  • 基于检索增强生成和分层调度策略,动态选择修复方法。
  • 在20,000对视频样本上训练,可区分7类退化与11种修复操作差异。
  • 适合复杂退化场景下的视频修复,尤其适用于真实世界应用。

真实场景中的视频修复面临异构退化问题,传统静态架构与固定推理流程难以泛化。近期基于智能体的方法虽具备动态决策能力,但受限于质量感知不足与搜索效率低下。本文提出VQ-Jarvis,一种融合检索增强与全链路智能的视频修复代理,具备更敏锐的视觉感知与更快的决策速度。为实现精准感知,构建了首个大规模视频成对增强数据集VSR-Compare,包含20,000对样本,覆盖7类退化类型、11种增强算子及多样内容域。基于该数据集,训练多算子判别模型与退化感知模型以指导决策。为提升效率,设计分层算子调度策略:对简单视频采用一步检索式路径获取;对复杂视频则执行逐步贪心搜索,在精度与效率间取得平衡。大量实验表明,VQ-Jarvis在复杂退化视频上持续优于现有方法。

原文摘要 · Abstract (English)

Video restoration in real-world scenarios is challenged by heterogeneous degradations, where static architectures and fixed inference pipelines often fail to generalize. Recent agent-based approaches offer dynamic decision making, yet existing video restoration agents remain limited by insufficient quality perception and inefficient search strategies. We propose VQ-Jarvis, a retrieval-augmented, all-in-one intelligent video restoration agent with sharper vision and faster thought. VQ-Jarvis is designed to accurately perceive degradations and subtle differences among paired restoration results, while efficiently discovering optimal restoration trajectories. To enable sharp vision, we construct VSR-Compare, the first large-scale video paired enhancement dataset with 20K comparison pairs covering 7 degradation types, 11 enhancement operators, and diverse content domains. Based on this dataset, we train a multiple operator judge model and a degradation perception model to guide agent decisions. To achieve fast thought, we introduce a hierarchical operator scheduling strategy that adapts to video difficulty: for easy cases, optimal restoration trajectories are retrieved in a one-step manner from a retrieval-augmented generation (RAG) library; for harder cases, a step-by-step greedy search is performed to balance efficiency and accuracy. Extensive experiments demonstrate that VQ-Jarvis consistently outperforms existing methods on complex degraded videos.

视频修复智能体检索增强动态决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。