arXiv:2512.07747cs.CV2025-12

Unison实现多模态任务全自动理解与生成,低成本高效通用。

Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation

  • 采用两阶段框架,融合预训练模型能力,低资源训练
  • 仅用50万样本和50小时GPU,覆盖理解与生成多种任务
  • 自动解析任务类型与元信息,无需人工配置

统一理解与生成是多模态学习中的重要方向。现有方法分为自回归训练和两阶段对齐微调两种,前者需海量数据与算力,后者虽成本较低但任务覆盖有限或生成质量差。两者均无法自动解析输入的元信息(如任务类型、图像分辨率、视频时长等),且依赖人工参数配置。本文提出Unison,采用两阶段方案,充分保留预训练模型能力。在极低训练成本下,覆盖文本、图像、视频理解,以及文生图、编辑、可控生成、基于知识产权的参考生成等多种生成任务。模型可自动识别用户意图,判断任务类型,并精准提取所需元信息,实现全流程自动化。实验表明,在仅50万训练样本和50 GPU小时的低成本设置下,模型能准确识别任务并提取参数,且在各类理解和生成任务中表现优异。

原文摘要 · Abstract (English)

Unified understanding and generation is a highly appealing research direction in multimodal learning. There exist two approaches: one trains a transformer via an auto-regressive paradigm, and the other adopts a two-stage scheme connecting pre-trained understanding and generative models for alignment fine-tuning. The former demands massive data and computing resources unaffordable for ordinary researchers. Though the latter requires a lower training cost, existing works often suffer from limited task coverage or poor generation quality. Both approaches lack the ability to parse input meta-information (such as task type, image resolution, video duration, etc.) and require manual parameter configuration that is tedious and non-intelligent. In this paper, we propose Unison which adopts the two-stage scheme while preserving the capabilities of the pre-trained models well. With an extremely low training cost, we cover a variety of multimodal understanding tasks, including text, image, and video understanding, as well as diverse generation tasks, such as text-to-visual content generation, editing, controllable generation, and IP-based reference generation. We also equip our model with the ability to automatically parse user intentions, determine the target task type, and accurately extract the meta-information required for the corresponding task. This enables full automation of various multimodal tasks without human intervention. Experiments demonstrate that, under a low-cost setting of only 500k training samples and 50 GPU hours, our model can accurately and automatically identify tasks and extract relevant parameters, and achieve superior performance across a variety of understanding and generation tasks.

多模态生成自动化低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。