arXiv:2602.12221cs.CV2026-02

统一离散流匹配框架,实现多模态理解与生成的高效协同。

Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching

  • 通过任务专用低秩适配器解耦理解与生成,避免目标冲突和表征纠缠。
  • 在八个基准上达到当前最优,零样本泛化能力覆盖图像修复等新任务。
  • 无需额外训练即可支持参考图像编辑、组合生成等复杂操作,适合实际应用。

我们提出UniDFlow,一种统一的离散流匹配框架,用于多模态理解、生成与编辑。该框架通过任务特定的低秩适配器解耦理解与生成过程,避免目标干扰与表征纠缠;同时引入基于参考的多模态偏好对齐机制,在相同条件下的相对结果优化中提升忠实度与可控性,且无需大规模重训练。UniDFlow在八个基准测试中取得当前最优性能,展现出强大的零样本泛化能力,可直接应用于图像修复、上下文图像生成、基于参考的编辑及组合生成等任务,尽管未进行任何显式任务训练。

原文摘要 · Abstract (English)

We propose UniDFlow, a unified discrete flow-matching framework for multimodal understanding, generation, and editing. It decouples understanding and generation via task-specific low-rank adapters, avoiding objective interference and representation entanglement, while a novel reference-based multimodal preference alignment optimizes relative outcomes under identical conditioning, improving faithfulness and controllability without large-scale retraining. UniDFlpw achieves SOTA performance across eight benchmarks and exhibits strong zero-shot generalization to tasks including inpainting, in-context image generation, reference-based editing, and compositional generation, despite no explicit task-specific training.

多模态生成模型流匹配零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。