arXiv:2511.08521cs.CV2025-11被引 11

UniVA让视频处理像搭积木一样,自动完成复杂任务

UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist

  • 用计划-执行双智能体架构,自动拆解用户指令为视频处理步骤
  • 支持多轮编辑、条件生成等复杂流程,可追溯每一步操作
  • 适合想构建交互式视频系统的开发者和研究者

尽管专用AI模型在视频生成或理解等单一任务上表现优异,但现实应用需要结合多种能力的复杂、迭代式工作流。为此,我们提出UniVA,一个开源的通用多智能体框架,可统一视频理解、分割、编辑与生成,形成连贯的工作流。UniVA采用计划-执行双智能体架构:规划器解析用户意图并分解为结构化步骤,执行器通过基于MCP的模块化工具服务器(分析、生成、编辑、追踪等)执行任务。借助分层多级记忆(全局知识、任务上下文、用户偏好),UniVA实现长时推理、上下文连续性与智能体间通信,支持交互式、自省式视频创作且全程可追溯。该设计实现以往需多个专用模型才能完成的复杂流程(如文本/图像/视频条件生成→多轮编辑→对象分割→组合合成)。我们还推出UniVA-Bench,一套涵盖理解、编辑、分割与生成的多步视频任务基准测试套件,用于严谨评估此类代理式视频系统。UniVA与UniVA-Bench均完全开源,旨在推动下一代多模态AI系统中交互式、代理式、通用型视频智能的研究。(https://univa.online/)

原文摘要 · Abstract (English)

While specialized AI models excel at isolated video tasks like generation or understanding, real-world applications demand complex, iterative workflows that combine these capabilities. To bridge this gap, we introduce UniVA, an open-source, omni-capable multi-agent framework for next-generation video generalists that unifies video understanding, segmentation, editing, and generation into cohesive workflows. UniVA employs a Plan-and-Act dual-agent architecture that drives a highly automated and proactive workflow: a planner agent interprets user intentions and decomposes them into structured video-processing steps, while executor agents execute these through modular, MCP-based tool servers (for analysis, generation, editing, tracking, etc.). Through a hierarchical multi-level memory (global knowledge, task context, and user-specific preferences), UniVA sustains long-horizon reasoning, contextual continuity, and inter-agent communication, enabling interactive and self-reflective video creation with full traceability. This design enables iterative and any-conditioned video workflows (e.g., text/image/video-conditioned generation $\rightarrow$ multi-round editing $\rightarrow$ object segmentation $\rightarrow$ compositional synthesis) that were previously cumbersome to achieve with single-purpose models or monolithic video-language models. We also introduce UniVA-Bench, a benchmark suite of multi-step video tasks spanning understanding, editing, segmentation, and generation, to rigorously evaluate such agentic video systems. Both UniVA and UniVA-Bench are fully open-sourced, aiming to catalyze research on interactive, agentic, and general-purpose video intelligence for the next generation of multimodal AI systems. (https://univa.online/)

视频代理多智能体视频生成开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。