arXiv:2603.28088cs.CV2026-03被引 9

让小模型通过记忆与技能实现复杂多模态生成

GEMS: Agent-Native Multimodal Generation with Memory and Skills

  • 构建多智能体闭环系统,持续优化生成质量
  • 用分层记忆减少冗余,全局追踪生成过程
  • 可加载专业技能,适配多样下游任务

近期多模态生成模型在通用任务上取得显著进展,但在复杂指令和特定下游任务上仍表现不足。受 Claude Code 等先进智能体框架启发,我们提出 GEMS(Agent-Native Multimodal GEneration with Memory and Skills),突破基础模型在通用与下游任务上的固有局限。GEMS 基于三大核心组件:Agent Loop 构建结构化多智能体框架,通过闭环优化迭代提升生成质量;Agent Memory 提供持久化的轨迹级记忆,分层存储事实状态与压缩的经验摘要,实现全局优化视图并降低冗余;Agent Skill 提供可扩展的领域专属能力集,支持按需加载,有效应对多样化下游应用。在五个主流任务和四个下游任务上,使用多个生成后端评估,GEMS 均实现显著性能提升。尤其值得注意的是,轻量级 6B 模型 Z-Image-Turbo 在 GenEval2 上超越了当前最先进模型 Nano Banana 2,验证了智能体驱动对模型能力的拓展效果。

原文摘要 · Abstract (English)

Recent multimodal generation models have achieved remarkable progress on general-purpose generation tasks, yet continue to struggle with complex instructions and specialized downstream tasks. Inspired by the success of advanced agent frameworks such as Claude Code, we propose \textbf{GEMS} (Agent-Native Multimodal \textbf{GE}neration with \textbf{M}emory and \textbf{S}kills), a framework that pushes beyond the inherent limitations of foundational models on both general and downstream tasks. GEMS is built upon three core components. Agent Loop introduces a structured multi-agent framework that iteratively improves generation quality through closed-loop optimization. Agent Memory provides a persistent, trajectory-level memory that hierarchically stores both factual states and compressed experiential summaries, enabling a global view of the optimization process while reducing redundancy. Agent Skill offers an extensible collection of domain-specific expertise with on-demand loading, allowing the system to effectively handle diverse downstream applications. Across five mainstream tasks and four downstream tasks, evaluated on multiple generative backends, GEMS consistently achieves significant performance gains. Most notably, it enables the lightweight 6B model Z-Image-Turbo to surpass the state-of-the-art Nano Banana 2 on GenEval2, demonstrating the effectiveness of agent harness in extending model capabilities beyond their original limits.

多模态生成智能体框架记忆机制小模型增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。