arXiv:2505.17613cs.AIcs.CL2025-05被引 4

构建了多模态生成评估新基准,显著提升自动化评价与人工判断的一致性。

MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation

  • 设计49项任务,覆盖4类模态组合,通过模型与程序协同实现可靠评估。
  • 人类评估一致性达94.3%,验证了基准的可靠性;最先进模型图像生成准确率仅78.3%。
  • 聚焦推理与交错生成难点,适合研究多模态生成模型能力边界的研究者使用。

自动评估多模态生成仍面临重大挑战,因自动化指标常难以与人工评价对齐,尤其在涉及多种模态的复杂任务中。为此,我们提出MMMG,一个全面且与人工评价高度一致的多模态生成基准,涵盖4种模态组合(图像、音频、文本与图像交织、文本与音频交织),聚焦对生成模型构成显著挑战的任务,同时通过模型与程序结合实现可靠的自动评估。MMMG包含49个任务(其中29个为新开发),每项任务配有精心设计的评估流程,共包含937条指令,系统评估模型的推理能力、可控性等关键性能。大量验证表明,MMMG与人工评价高度一致,平均吻合度达94.3%。在24个多模态生成模型上的基准测试显示,尽管最先进模型GPT Image在图像生成上达到78.3%准确率,但在多模态推理和交错生成任务上表现不足。结果还表明音频生成仍有巨大提升空间,指明未来研究的重要方向。

原文摘要 · Abstract (English)

Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align reliably with human evaluation, especially for complex tasks that involve multiple modalities. To address this, we present MMMG, a comprehensive and human-aligned benchmark for multimodal generation across 4 modality combinations (image, audio, interleaved text and image, interleaved text and audio), with a focus on tasks that present significant challenges for generation models, while still enabling reliable automatic evaluation through a combination of models and programs. MMMG encompasses 49 tasks (including 29 newly developed ones), each with a carefully designed evaluation pipeline, and 937 instructions to systematically assess reasoning, controllability, and other key capabilities of multimodal generation models. Extensive validation demonstrates that MMMG is highly aligned with human evaluation, achieving an average agreement of 94.3%. Benchmarking results on 24 multimodal generation models reveal that even though the state-of-the-art model, GPT Image, achieves 78.3% accuracy for image generation, it falls short on multimodal reasoning and interleaved generation. Furthermore, results suggest considerable headroom for improvement in audio generation, highlighting an important direction for future research.

多模态生成评估基准人机对齐模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。