评测统一多模态模型在生成与理解任务间的语义一致性
Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

- 基于场景图构建生成与理解任务的对齐测试
- 发现高精度不等于跨任务语义一致
- 适合关注模型内在一致性而非单一性能的研究者
统一多模态模型(uMMs)旨在共享表征以同时支持视觉理解和生成。然而,现有评估方法独立考察这两项能力,未检验其语义是否对齐。为此,我们提出XTC-Bench,一个基于场景图的评估框架,通过从结构化场景图中提取生成提示和理解问题,实现对物体、属性和关系等原子事实的细粒度语义对齐分析。我们提出连续跨任务一致度(CCTA),量化生成与理解在匹配事实上的语义一致性,分离出内部一致性与单任务准确率。在八个开源和一个商用统一模型上的实验表明,高生成或理解性能并不意味着强跨任务对齐;架构分析显示,一致性取决于跨模态学习目标的耦合紧密度,而非仅由架构统一决定。XTC-Bench提供可复现、模型无关的诊断框架,推动多模态建模超越孤立任务表现。
原文摘要 · Abstract (English)
Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evaluation protocols assess these two capabilities independently and do not examine whether they are semantically aligned. As a result, it remains unclear whether current uMMs learn coherent unified representations that remain consistent across tasks given a visual concept. We introduce XTC-Bench, a scene-graph-grounded evaluation framework that measures cross-task visual semantic consistency. By deriving both generation prompts and understanding queries from a structured scene graph, our framework enables fact-level alignment analysis across objects, attributes, and relations. We propose Continuous Cross-Task Agreement (CCTA), a fine-grained metric that quantifies semantic agreement between generation and understanding over matched atomic facts, isolating internal consistency from standalone task accuracy. Extensive experiments on eight open-source and one commercial unified models reveal that high generation or understanding performance does not imply strong cross-task alignment, and architectural analysis shows consistency is governed by how tightly learning objectives are coupled across modalities, not by architectural unification alone. XTC-Bench provides a reproducible and model-agnostic framework for diagnosing representation-level misalignment, offering a concrete direction for advancing unified multimodal modeling beyond isolated task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。