用符号语言统一描述任意生成任务,无需训练即可执行
Symbolic Representation for Any-to-Any Generative Tasks
- 用函数、参数和拓扑逻辑构建符号化任务流程
- 12种跨模态生成任务均零训练完成,性能媲美顶尖模型
- 支持自由编辑与中断,适合快速原型开发
我们提出一种符号化生成任务描述语言及对应推理引擎,可将任意多模态生成任务表示为结构化的符号流。不同于依赖大规模训练和隐式神经表示的传统生成模型,该框架采用显式符号表示,包含函数、参数和拓扑逻辑三大核心组件。借助预训练语言模型,推理引擎可无训练地将自然语言指令直接映射为符号工作流。实验表明,该方法成功完成超过12种多样化的多模态生成任务,在内容质量上达到或超越现有最先进统一模型水平,同时具备更高效率、更强可编辑性和可中断性。我们认为,符号化任务表示为生成式AI的低成本、可扩展发展提供了坚实基础。
原文摘要 · Abstract (English)
We propose a symbolic generative task description language and a corresponding inference engine capable of representing arbitrary multimodal tasks as structured symbolic flows. Unlike conventional generative models that rely on large-scale training and implicit neural representations to learn cross-modal mappings, often at high computational cost and with limited flexibility, our framework introduces an explicit symbolic representation comprising three core primitives: functions, parameters, and topological logic. Leveraging a pre-trained language model, our inference engine maps natural language instructions directly to symbolic workflows in a training-free manner. Our framework successfully performs over 12 diverse multimodal generative tasks, demonstrating strong performance and flexibility without the need for task-specific tuning. Experiments show that our method not only matches or outperforms existing state-of-the-art unified models in content quality, but also offers greater efficiency, editability, and interruptibility. We believe that symbolic task representations provide a cost-effective and extensible foundation for advancing the capabilities of generative AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。