arXiv:2602.08336cs.CLcs.CV2026-02被引 5

测试统一多模态模型中推理与生成的对齐程度,发现推理反而拖后腿。

UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models

  • 用推理引导图像生成作为诊断任务,对比三种生成方式。
  • 去上下文生成比推理引导生成效果更好,提升达15%以上。
  • 适合研究多模态对齐、模型设计或评估的学者使用。

统一多模态模型(UMMs)旨在将多模态理解与生成整合于同一架构中,但其文本与视觉模态的对齐程度尚不明确。为此,我们提出以推理引导图像生成为诊断任务,模型先生成文本推理,再据此生成图像。构建了包含2,000个经人工标注与验证实例的UReason基准,覆盖代码、算术、空间、属性和文本五类高推理强度任务。设计评估框架,对比直接生成、推理引导生成及仅依赖推理提炼提示的去上下文生成。在八款主流开源UMMs上,尽管推理引导生成优于直接生成,但令人意外的是,去上下文生成始终显著领先,性能提升超过15%。进一步分析表明,文本推理中的预期视觉语义未能可靠映射至生成图像中,即便模型采用统一架构与训练。UReason可作为衡量推理-生成对齐的有效指标,为下一代更紧密对齐的多模态模型提供挑战性基准。

原文摘要 · Abstract (English)

Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent textual and visual modalities are aligned. To investigate this question, we use reasoning-guided image generation as a diagnostic task, where models produce textual reasoning first and then generate images. We introduce UReason, a benchmark for evaluating reasoning-to-generation alignment in this paradigm, consisting of 2,000 human-curated and human-verified instances spanning five reasoning-intensive tasks: Code, Arithmetic, Spatial, Attribute, and Text. To enable controlled analysis, we develop an evaluation framework that compares direct generation, reasoning-guided generation, and decontextualized generation, which conditions only on the refined prompt extracted from reasoning. Across eight widely used open-source UMMs, while we find that reasoning-guided generation yields improvements over direct generation, somewhat surprisingly, decontextualized generation consistently outperforms reasoning-guided generation by a large margin. Our further analyses suggest that the intended visual semantics in textual reasoning are not reliably reflected in the generated images, despite their unified design and training. Overall, UReason serves as a practical litmus test for reasoning-to-generation alignment and provides a challenging benchmark for developing next-generation, more tightly aligned UMMs.

多模态推理对齐生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。