arXiv:2511.01163cs.CV2025-11被引 12

评测模型在图文互推中的跨模态推理能力,发现双向交互决定生成质量。

ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation

  • 设计双场景人类标注数据集,测试图文互推推理
  • 交叉模态推理显著提升图像生成质量,非交替模型表现差
  • 模型能理解直观概念但难处理抽象符号,适合研究多模态智能

统一多模态模型(UMMs)已成为融合文本与图像理解与生成的有力范式。然而,现有评估方法将这些能力孤立看待:多模态输入输出任务主要通过单模态推理评分,即文本基准侧重语言推理,视觉基准侧重像素层面的推理结果。为此,我们提出ROVER,旨在评测核心的双向跨模态推理能力——即利用一种模态指导、验证或优化另一种模态的输出。ROVER是一个人工标注的基准,明确针对双向跨模态推理,包含1312个任务,基于1876张图像,涵盖两个互补场景:(1)以语言增强的视觉生成推理,测试模型能否通过语言提示和推理链准确生成图像;(2)以视觉增强的语言生成推理,测试模型能否生成中间可视化辅助自身问答推理。在17个统一模型上的实验揭示两项关键发现:(i)跨模态推理决定视觉生成质量,交替处理模型显著优于非交替模型;即使强单模态模型组合也无法达到同等推理水平。(ii)模型在物理与符号推理间存在分离:能准确解读感知概念,但在符号任务中难以构建视觉抽象,错误推理严重损害性能。结果表明,双向跨模态推理是实现真正全模态生成的关键前沿。

原文摘要 · Abstract (English)

Unified multimodal models (UMMs) have emerged as a powerful paradigm for seamlessly unifying text and image understanding and generation. However, prevailing evaluations treat these abilities in isolation, such that tasks with multimodal inputs and outputs are scored primarily through unimodal reasoning, i.e., textual benchmarks emphasize language-based reasoning, while visual benchmarks emphasize reasoning outcomes manifested in the pixels. We introduce ROVER to address this pressing need to test reciprocal cross-modal reasoning, the use of one modality to guide, verify, or refine outputs in the other, an ability central to the vision of unified multimodal intelligence. ROVER is a human-annotated benchmark that explicitly targets reciprocal cross-modal reasoning, which contains 1312 tasks grounded in 1876 images, spanning two complementary settings. Verbally-augmented reasoning for visual generation evaluates whether models can use verbal prompts and reasoning chains to guide faithful image synthesis. Visually-augmented reasoning for verbal generation evaluates whether models can generate intermediate visualizations that strengthen their own reasoning processes for question answering. Experiments on 17 unified models reveal two key findings: (i) Cross-modal reasoning determines visual generation quality, with interleaved models significantly outperforming non-interleaved ones; notably, combining strong unimodal models fails to achieve comparable reasoning. (ii) Models show dissociation between physical and symbolic reasoning: they succeed at interpreting perceptual concepts literally but fail to construct visual abstractions for symbolic tasks, where faulty reasoning harms performance. These results highlight reciprocal cross-modal reasoning as a critical frontier for enabling true omnimodal generation.

跨模态推理多模态生成图文互推评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。