arXiv:2606.27377cs.CVcs.CL2026-06被引 7

让生成模型同时学会作图、局部和全局编辑而不互相干扰

DanceOPD: On-Policy Generative Field Distillation

  • 用动态路由把每张图分到对应能力领域,统一训练
  • 在保留原图质量基础上,多任务表现显著提升
  • 适合需要多功能生成的模型开发与部署

现代图像生成要求单一模型具备文本生图(T2I)、局部编辑和全局编辑等多种能力,但这些能力常不兼容且相互冲突。例如,编辑会降低文本生图性能,而局部与全局编辑也会相互干扰。为此,我们提出DanceOPD——一种面向流匹配模型的在线策略生成场蒸馏框架。该方法将每个样本路由至特定能力场,查询低噪声的学生状态,并以简单的速度均方误差目标进行训练。通过将各能力源定义为共享流状态空间上的速度场,学生模型在自身推演状态上学习来自各专家场的知识,实现能力融合。该框架还能吸收分类器自由引导等外部控制场。在文本生图、编辑、真实感场吸收及CFG吸收等多个任务上的综合实验表明,本方法有效提升了多能力组合性能,在增强目标能力的同时保持了原始生成质量。我们认为该工作为流匹配模型中的生成场蒸馏提供了实用路径。

原文摘要 · Abstract (English)

Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are rarely naturally aligned and often conflict. For instance, editing tends to degrade T2I performance, while global and local editing interfere with each other. Consequently, effectively composing these capabilities has become a central challenge for image generation model training. To tackle this, we introduce DanceOPD, an on-policy generative field distillation framework for flow-matching models that routes each sample to one capability field, queries one low-noise student-induced state, and trains with a simple velocity MSE objective. With each capability source defined as a velocity field over the shared flow state space, the student learns from fields queried on its own rollout states to compose expert capabilities. This formulation also absorbs operator-defined fields such as classifier-free guidance. Comprehensive experiments on T2I, editing, realism-field absorption, and CFG absorption show that our approach improves multi-capability composition, strengthening target capabilities while preserving anchor generation quality. We believe this work establishes a practical route for generative field distillation in flow-matching models.

生成模型多任务生成流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。