arXiv:2603.05075cs.CV2026-03被引 7

首个统一任意到任意交叉模态评测基准,推动多模态大模型理解生成能力

UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark

  • 构建跨30领域、7模态的任意交错输入输出数据集
  • 涵盖31000个高质量样本,测试语义正确性与交错连贯性
  • 适合研究多模态推理与生成的学者及工业开发者

现实中的多模态应用需处理用户任意组合且交错的多模态输入,并生成任意交错的多媒体输出。这定义了统一范式下任意到任意交错多模态学习的目标,为多模态大语言模型(MLLMs)带来新挑战与机遇。本文提出UniM基准,首个统一的任意到任意交错多模态数据集,包含31,000个高质量实例,覆盖30个领域和7种代表性模态:文本、图像、音频、视频、文档、代码、3D。我们进一步提出UniM评估套件,从语义正确性与生成质量、响应结构完整性、交错连贯性三个维度评估模型。此外,我们设计了UniMA代理基线模型,具备可追溯推理能力以实现结构化交错生成。全面实验揭示了UniM的难度,指出了推进统一多模态智能的关键挑战与方向。项目主页:https://any2any-mllm.github.io/unim。

原文摘要 · Abstract (English)

In real-world multimodal applications, systems usually need to comprehend arbitrarily combined and interleaved multimodal inputs from users, while also generating outputs in any interleaved multimedia form. This capability defines the goal of any-to-any interleaved multimodal learning under a unified paradigm of understanding and generation, posing new challenges and opportunities for advancing Multimodal Large Language Models (MLLMs). To foster and benchmark this capability, this paper introduces the UniM benchmark, the first Unified Any-to-Any Interleaved Multimodal dataset. UniM contains 31K high-quality instances across 30 domains and 7 representative modalities: text, image, audio, video, document, code, and 3D, each requiring multiple intertwined reasoning and generation capabilities. We further introduce the UniM Evaluation Suite, which assesses models along three dimensions: Semantic Correctness & Generation Quality, Response Structure Integrity, and Interleaved Coherence. In addition, we propose UniMA, an agentic baseline model equipped with traceable reasoning for structured interleaved generation. Comprehensive experiments demonstrate the difficulty of UniM and highlight key challenges and directions for advancing unified any-to-any multimodal intelligence. The project page is https://any2any-mllm.github.io/unim.

多模态评测基准生成推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。