评测大模型跨模态能力的新基准,覆盖五种模态组合。
Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models

- 构建涵盖文本、图像、音频等五类模态的多模态评估基准
- 五款前沿模型在模态生成上表现有限,最高仅34.9分
- 提出模态存在率指标,更精准衡量模型输出完整性
前沿语言模型日益被宣传为可跨模态感知与响应的全模态系统。现有评估框架却几乎仅聚焦双模态理解(通常为文本加一种其他模态)。我们提出模态成熟度指数(MMI),一个用于评估大语言模型在文本、图像、音频、视频和文档五种模态及最多三种模态组合的多模态能力的基准。MMI包含893个精心设计的问题,每个问题要求模型展示对多重输入模态的理解,并生成包含多种输出格式的回应。问题自包含,对正确响应所需模态有明确要求。每条提示附有人工编写的评分标准,模型的MMI值为其在各模态上的平均得分。由于低分可能源于模态缺失或内容错误,我们引入补充指标模态存在率(MPS),即对预期输出模态的每题F1分数。对五款前沿多模态模型的应用显示,其MPS介于15.6(Claude Opus 4.6)至34.9(GPT-5.4)之间。鉴于生成模态极少,报告以MPS为主结果,待模型改进。为评估基于LLM判断者与评分标准的有效性,我们进行独立实验使用定制生成工具,在生成资产上发现,采用评分标准的LLM判断者与未见标准的人类标注者在70.8%的判断上达成一致。
原文摘要 · Abstract (English)
Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities. Existing evaluation frameworks, however, focus almost exclusively on bimodal understanding, typically text plus one other modality. We propose the Modality Maturity Index (MMI), a benchmark designed to evaluate the multimodal capabilities of large language models across five modalities (text, image, audio, video and document) and combinations of up to three modalities in both inputs and outputs. MMI consists of 893 questions, each carefully crafted to require the model to demonstrate its understanding of multiple input modalities and to generate responses that incorporate various output formats. The questions are designed to be self-contained, with clear expectations for the correct modality or mix of modalities required for an accurate response. Every MMI prompt carries human-authored rubric criteria for each output modality expected in the response; a model's MMI Value expresses the average of the per-modality scores for each prompt. Because low scores can reflect either failure to generate a modality (lack of presence) or failure to generate correct content, we introduce also a supplementary Modality Presence Score (MPS), a per-prompt F1 over the expected output modalities. Applying MMI to five frontier multimodal models, we find that the MPS ranges from only 15.6 (Claude Opus 4.6) to 34.9 (GPT-5.4). Given the low availability of returned modalities to even grade, we report MPS as our main result pending model improvements. To assess the viability of judging output correctness with LLM judges and rubrics, we run a separate experiment with custom generation tools. On the assets that generates, we find that an LLM judge applying the rubrics agrees with rubric-blind human annotators (who score the outputs directly and never see the criteria) on 70.8% of judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。