arXiv:2602.06442cs.CV2026-02被引 2

让多模态模型能持续对话,流畅处理图文交错的复杂交互。

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation

  • 通过连续对话训练和数据合成,让模型理解跨轮次上下文。
  • 在视觉理解与指令编辑任务上超越开源模型,保持图像生成质量。
  • 适合需要长对话、多轮图文交互的应用场景。

统一多模态模型(UMMs)虽取得显著进展,但仍受限于单轮交互范式,更像是独立请求的解决者而非持续对话助手。为此,我们提出 ChatUMM——一种具备鲁棒上下文追踪能力的对话式统一模型,可支持交错的多模态生成。其核心创新包括:一种将文本-图像序列建模为连续对话流的交错多轮训练策略,以及一套系统化的对话数据合成流程。该流程分三阶段将多样单轮数据转化为自然对话:构建带状态的基础对话、通过含历史依赖重写的干扰轮次强化长程依赖推理、合成自然交错的多模态响应。大量实验表明,ChatUMM 在视觉理解与指令引导编辑基准上达到开源统一模型最优表现,同时在文本到图像生成中保持竞争力。尤其在复杂多轮场景中表现出更强鲁棒性,确保对话连贯且上下文感知。

原文摘要 · Abstract (English)

Unified multimodal models (UMMs) have achieved remarkable progress yet remain constrained by a single-turn interaction paradigm, effectively functioning as solvers for independent requests rather than assistants in continuous dialogue. To bridge this gap, we present ChatUMM. As a conversational unified model, it excels at robust context tracking to sustain interleaved multimodal generation. ChatUMM derives its capabilities from two key innovations: an interleaved multi-turn training strategy that models serialized text-image streams as a continuous conversational flow, and a systematic conversational data synthesis pipeline. This pipeline transforms a diverse set of standard single-turn datasets into fluid dialogues through three progressive stages: constructing basic stateful dialogues, enforcing long-range dependency resolution via ``distractor'' turns with history-dependent query rewriting, and synthesizing naturally interleaved multimodal responses. Extensive evaluations demonstrate that ChatUMM achieves state-of-the-art performance among open-source unified models on visual understanding and instruction-guided editing benchmarks, while maintaining competitive fidelity in text-to-image generation. Notably, ChatUMM exhibits superior robustness in complex multi-turn scenarios, ensuring fluid, context-aware dialogues.

多模态对话上下文追踪生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。