SeedEdit 3.0实现快速高质量图像编辑,保真度与指令遵循能力显著提升。
SeedEdit 3.0: Fast and High-Quality Generative Image Editing
- 采用元信息管道与嵌入策略,融合多源数据增强训练
- 联合学习扩散损失与奖励损失,提升编辑一致性
- 真实图像编辑可用率达56.1%,优于前代及主流模型
我们提出SeedEdit 3.0,搭配T2I模型Seedream 3.0,显著提升在真实图像输入上的指令遵循能力与内容(如身份、风格)保留效果。相比以往版本,本报告引入三项关键改进:首先,构建基于元信息范式与嵌入策略的增强数据整理流程,支持跨数据源图像混合,有效扩展编辑数据规模,并促进视觉语言模型与扩散模型更紧密协同;其次,设计联合学习框架,同时优化扩散损失与奖励损失;最后,在测试基准上评估真实与合成图像编辑任务,结果表明其在多维度间取得最佳平衡,可用率达56.1%,显著优于SeedEdit 1.6(38.4%)、GPT4o(37.1%)和Gemini 2.0(30.3%)。
原文摘要 · Abstract (English)
We introduce SeedEdit 3.0, in companion with our T2I model Seedream 3.0, which significantly improves over our previous SeedEdit versions in both aspects of edit instruction following and image content (e.g., ID/IP) preservation on real image inputs. Additional to model upgrading with T2I, in this report, we present several key improvements. First, we develop an enhanced data curation pipeline with a meta-info paradigm and meta-info embedding strategy that help mix images from multiple data sources. This allows us to scale editing data effectively, and meta information is helpfult to connect VLM with diffusion model more closely. Second, we introduce a joint learning pipeline for computing a diffusion loss and reward losses. Finally, we evaluate SeedEdit 3.0 on our testing benchmarks, for real/synthetic image editing, where it achieves a best trade-off between multiple aspects, yielding a high usability rate of 56.1%, compared to SeedEdit 1.6 (38.4%), GPT4o (37.1%) and Gemini 2.0 (30.3%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。