构建首个覆盖56类真实场景的图文交错生成评测基准
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

- 设计5400条人工标注数据,覆盖旅行、设计等多元任务
- 提出IntJudge模型,与人类判断一致率达82.42%,优于GPT评估器11.34%
- 揭示现有模型在图文交错生成上仍有巨大提升空间
多模态大语言模型在视觉理解与生成任务上取得显著进展,但图文交错内容生成仍具挑战,需整合多模态理解与生成能力。现有评测基准因数据量和多样性不足,难以有效评估此类方法。为此,我们提出OpenING,一个包含5,400条高质量人工标注实例、覆盖56个真实世界任务的综合性评测基准。其涵盖旅行指南、设计、头脑风暴等多样化日常场景,为挑战图文交错生成方法提供坚实平台。此外,我们提出IntJudge,一种用于评估开放式多模态生成的方法的判别模型。通过新型数据流水线训练,IntJudge与人类判断的一致率达到82.42%,优于基于GPT的评估器11.34%。在OpenING上的大量实验表明,当前图文交错生成方法仍有巨大改进空间。关键发现进一步为下一代模型的发展提供指导。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a challenge, which requires integrated multimodal understanding and generation abilities. While the progress in unified models offers new solutions, existing benchmarks are insufficient for evaluating these methods due to limitations in data size and diversity. To bridge this gap, we introduce OpenING, a comprehensive benchmark comprising 5,400 high-quality human-annotated instances across 56 real-world tasks. OpenING covers diverse daily scenarios such as travel guide, design, and brainstorming, offering a robust platform for challenging interleaved generation methods. In addition, we present IntJudge, a judge model for evaluating open-ended multimodal generation methods. Trained with a novel data pipeline, our IntJudge achieves an agreement rate of 82.42% with human judgments, outperforming GPT-based evaluators by 11.34%. Extensive experiments on OpenING reveal that current interleaved generation methods still have substantial room for improvement. Key findings on interleaved image-text generation are further presented to guide the development of next-generation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。