arXiv:2607.12042cs.CVcs.LG2026-07

让模型像人一样不断学习新概念,自动进化生成能力。

SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning

论文配图:SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning
图 1 · 摘自论文原文
  • 用符号化概念框抽象操作,实现可复用的知识积累
  • 比现有系统生成质量更高,任务成功率提升40%以上
  • 适合需要持续学习的智能创作场景,如交互式设计

视觉生成在文本到图像/视频合成、多模态交互创作等众多领域日益普及。然而,现有单体模型普遍存在无法累积学习和自主演化的缺陷,我们称之为“永久新手”问题。它们缺乏将经验结构化为可复用知识的机制,因此每次任务都依赖脆弱的“从零开始”推理,导致组合泛化能力差且知识保留效率低。为此,我们提出SymbOmni,一种通过符号概念学习实现累积进化的智能全能模型。其核心是可优化的符号概念盒,能将低层操作抽象为可重用的符号工作流指令。SymbOmni通过归纳-转导循环运行:先将经验抽象为符号概念(归纳),再自适应组合以解决新任务(转导)。训练采用语言反馈的显式反向传播,实现无需梯度微调的持续自我改进。全面实验表明:(I) SymbOmni在迭代创作任务中显著优于现有基于代理的系统,并超越封闭源代码模型(如Nano Banana、GPT-Image-1)在图像质量和任务成功率上的表现;(II) SymbOmni将令牌消耗降低超40%,同时保持竞争力生成质量;(III) SymbOmni实现了有效持续学习,在多个在线学习基准上达成累积收益,创下新基准。

原文摘要 · Abstract (English)

Visual generation is increasingly ubiquitous in diverse domains, from text-to-image/video synthesis to multimodal interactive creation. Yet prevailing monolithic models remain fundamentally constrained by their inability to learn cumulatively and evolve autonomously, which is a limitation we term the "perpetual novice" problem. They lack mechanisms for structuring experience into reusable knowledge and therefore rely on brittle, "from-scratch" reasoning for each task, resulting in poor compositional generalization and inefficient knowledge retention. Motivated by these limitations, we propose SymbOmni, an agentic omni-model designed for cumulative evolution through Symbolic Concept Learning. At its core is the Symbolic Concept Box, an optimizable memory module that abstracts low-level operations into reusable Symbolic Workflow Instructions. SymbOmni operates through an induction-transduction cycle: experiences are abstracted into symbolic concepts (induction), which are then adaptively composed to solve novel tasks (transduction). The training is done by verbalized backpropagation with language-based feedback to enable continuous self-improvement without gradient-based model fine-tuning. Comprehensive experiments validate that (I) SymbOmni significantly outperforms existing agent-based systems for iterative creation and also surpasses closed-source models (e.g., Nano Banana, GPT-Image-1) in both image quality and task success rates; (II) SymbOmni effectively reduces token consumption by over 40% while maintaining competitive generation quality; and (III) SymbOmni enables effective continual learning by achieving cumulative gains across multiple online-learning benchmarks and setting a new state of the art.

符号学习持续学习智能生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。