提出闭环评估框架,让多模态模型自检生成内容的逻辑一致性。
A Model-Internal Protocol for Assessing Multimodal Models as Integrated Systems

- 用自我生成-理解闭环测试模型整合能力
- 高分模型在自生内容上仍常推理失败
- 适合评估统一架构多模态模型的系统级表现
随着大视觉语言模型日益将视觉生成与理解整合于单一参数空间,如何以协同方式评估这种结构统一性仍是关键挑战。现有评估协议通常将生成与判别能力分开处理,未能覆盖统一多模态模型(UMMs)的系统级评估。本文提出无需标注的自我生成-理解(SGU)评估框架,通过语义闭环挑战探测统一模型的集成能力:要求模型先感知图像并生成描述,再基于该描述重建视觉上下文,最后对自生成输出进行推理。该流程无需新增标注,提供零成本测试平台,生成专用于评估统一系统性能的综合得分。大量实验表明,即使高性能的UMMs在自身生成内容上的推理也常出现困难,揭示了独立评估理解或生成无法捕捉的局限性。本工作为下一代统一多模态模型提供了互补的全貌评估框架。
原文摘要 · Abstract (English)
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluating such structural unification in a cohesive manner remains a critical challenge. Current evaluation protocols predominantly treat generative and discriminative capabilities as separate tasks, leaving a gap in system-level evaluation for unified multimodal models (UMMs). In this work, we propose Self-Generative-Understanding (SGU), a novel, annotation-free evaluation framework that probes the integrated capabilities of unified models through a semantic closed-loop challenge. Without requiring new annotations, SGU leverages the dual understanding-and-generation abilities of UMMs by asking them to first perceive an image and produce a textual description, subsequently reconstruct a visual context based on that description, and finally perform reasoning over the self-generated output. This pipeline provides a zero-cost testbed that yields an integrated performance score specifically tailored for evaluating UMMs as unified systems. Extensive experiments show that even high-performing UMMs often struggle to reason over their own generated contexts, revealing limitations that are not captured by separate evaluations of understanding or generation alone. Our work provides a complementary holistic evaluation framework and offers a foundation for benchmarking the development of next-generation unified multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。