arXiv:2601.00561cs.CV2026-01

构建多任务评测基准,揭示统一多模态模型的世界知识短板

AEGIS: Exploring the Limit of World Knowledge Capabilities for Unified Mulitmodal Models

  • 设计AEGIS多任务评估框架,覆盖视觉理解、生成与编辑等
  • 1050个手动标注问题显示多数模型在复杂推理中表现差
  • 提出确定性检查清单评估法,提升评测可靠性,适合研究者参考

统一多模态模型(UMMs)在跨任务应用世界知识的能力仍面临重大挑战。现有基准仅提供孤立的单任务评估,诊断能力有限。为此,我们提出AEGIS(Assessing Editing, Generation, Interpretation-Understanding for Super-intelligence),一个涵盖视觉理解、生成、编辑及混合生成的综合性多任务基准。AEGIS包含1,050个精心设计的手动标注问题,覆盖21个主题(包括STEM、人文学科、日常生活等)和6种推理类型。为准确评估模型的世界知识范围,避免模糊评分,我们进一步提出确定性检查清单评估(DCE),以原子级“是/否”判断替代传统提示式打分,显著提升评估可靠性。大量实验表明,多数UMMs存在严重世界知识缺陷,且复杂推理下性能大幅下降。此外,简单插入推理模块可部分缓解这些问题,揭示未来研究的重要方向。结果强调了基于世界知识的推理对UMMs发展的关键意义。

原文摘要 · Abstract (English)

The capability of Unified Multimodal Models (UMMs) to apply world knowledge across diverse tasks remains a critical, unresolved challenge. Existing benchmarks fall short, offering only siloed, single-task evaluations with limited diagnostic power. To bridge this gap, we propose AEGIS (\emph{i.e.}, \textbf{A}ssessing \textbf{E}diting, \textbf{G}eneration, \textbf{I}nterpretation-Understanding for \textbf{S}uper-intelligence), a comprehensive multi-task benchmark covering visual understanding, generation, editing, and interleaved generation. AEGIS comprises 1,050 challenging, manually-annotated questions spanning 21 topics (including STEM, humanities, daily life, etc.) and 6 reasoning types. To concretely evaluate the performance of UMMs in world knowledge scope without ambiguous metrics, we further propose Deterministic Checklist-based Evaluation (DCE), a protocol that replaces ambiguous prompt-based scoring with atomic ``Y/N'' judgments, to enhance evaluation reliability. Our extensive experiments reveal that most UMMs exhibit severe world knowledge deficits and that performance degrades significantly with complex reasoning. Additionally, simple plug-in reasoning modules can partially mitigate these vulnerabilities, highlighting a promising direction for future research. These results highlight the importance of world-knowledge-based reasoning as a critical frontier for UMMs.

多模态模型世界知识评测基准推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。