arXiv:2607.23588cs.CV2026-07

构建可编辑画布上的多模态创作代理系统,支持长期协作创作。

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

论文配图:JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
图 1 · 摘自论文原文
  • 以可编辑画布作为代理的外部记忆和共享工作区,统一管理创作状态。
  • 通过三层架构实现代理在可追溯、可干预的创作环境中持续生成与修订。
  • 适合研究长周期多模态创作的自动化与人机协同机制。

创意AI正从单步内容生成转向长周期多模态创作。尽管当前生成模型能高质量合成图像、视频、音频、UI元素、故事板、幻灯片等创意资产,但真实创作远不止孤立的提示-输出交互。它包含参考、草稿、备选方案、修改、失败尝试、版本关系、工具操作、评估信号和人类反馈,共同构成动态演进的项目状态。现有基于提示、对话或节点的生成系统仅部分支持此状态,常丢失中间上下文,依赖线性对话,或需手动定义流程。近期商业系统显示向代理辅助创作转变,但其封闭架构使研究代理如何表征上下文、选择工具、修正作品、从失败中恢复及保持长期一致性变得困难。为此,我们提出JarvisHub——一个面向长周期多模态创作的原生画布式创作代理框架。JarvisHub将可编辑画布作为用户工作区、代理外部记忆、动作空间与共享项目状态,以类型化的画布节点和链接表示多模态资产、依赖关系、版本和反馈。通过画布状态、协议桥接与代理运行时三层架构,实现代理在可检查、可编辑的创作状态中行动。该设计使创作代理超越孤立工具使用,迈向可持续、可人工引导的创作自动化,支持代理在用户全程可监控、可指导、可介入下,逐步规划、生成、修订并组织多模态项目。

原文摘要 · Abstract (English)

Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state. Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift toward agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time. To address this gap, we introduce JarvisHub, a canvas-native creative agent harness for long-horizon multimodal creation. JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use toward sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.

多模态生成创作代理人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。