arXiv:2605.21487cs.CV2026-05被引 3

用智能编辑统一训练多模态模型,三能力同步提升。

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning

论文配图:Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning
图 1 · 摘自论文原文
  • 以智能编辑为单一任务,整合理解与生成能力
  • 仅用一个数据集和训练阶段,实现三能力全面增强
  • 适合想简化多任务训练的AI研发者

当前提升统一多模态模型(UMMs)在图像理解、生成与编辑方面的能力,主要依赖混合多任务训练。由于任务间存在固有冲突,该策略需复杂的多阶段流程、大量数据混合及平衡技巧,最终仅带来性能权衡而非真正协同增益。为此,我们提出Uni-Edit,一种作为统一模型调优通用任务的智能图像编辑任务。不同于复杂混合流程,Uni-Edit仅通过一个任务、一次训练阶段和一个数据集,同时提升三种能力。我们首先识别图像编辑作为理想通用任务的潜力,因其天然融合视觉理解与生成需求。然而,现有编辑数据依赖简单指令,严重低估模型理解能力。为此,我们提出首个自动化可扩展的数据合成流水线,将多样化VQA数据转化为包含嵌入问题与嵌套逻辑的复杂编辑指令。由此生成Uni-Edit-148k数据集,配对多样推理密集型指令与高质量编辑图像。在BAGEL和Janus-Pro上的实验表明,仅在Uni-Edit上进行调训即可实现三能力全面增强,无需任何辅助操作。

原文摘要 · Abstract (English)

Currently, enhancing Unified Multimodal Models (UMMs) with image understanding, generation, and editing capabilities mainly relies on mixed multi-task training. Due to inherent task conflicts, such strategy requires complex multi-stage pipelines, massive data mixing, and balancing tricks, merely resulting in a performance trade-off rather than true mutual reinforcement. To break this paradigm, we propose Uni-Edit, an intelligent image editing task that serves as the first general task for UMM tuning. Unlike complex mixed pipelines, Uni-Edit improves performance across all three abilities at once using only one task, one training stage, and one dataset. Specifically, we first identify image editing as an inherently ideal general task, as it naturally demands both visual understanding and generation. However, existing editing data relies on simplistic instructions that severely underutilize a model's understanding capacity. To address this, we introduce the first automated and scalable data synthesis pipeline for intelligent editing, transforming diverse VQA data into complex and effective editing instructions with embedded questions and nested logic. This yields Uni-Edit-148k, pairing diverse reasoning-intensive instructions with high-quality edited images. Extensive experiments on BAGEL and Janus-Pro demonstrate that tuning solely on Uni-Edit achieves comprehensive enhancements across all three capabilities without any auxiliary operations.

多模态图像编辑统一模型数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。