arXiv:2510.00069cs.CV2025-10

构建多智能体标注的图文指南理解基准,评测大模型对复杂图文关系的理解能力。

OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding

  • 用多智能体协作生成图像描述,降低人工标注成本
  • 29个主流多模态模型平均准确率77%,最优为Qwen2.5-VL-72B
  • 揭示当前模型在语义理解与逻辑推理上的明显短板

近年来多模态大语言模型(MLLMs)展现出强大能力,但对其在图文指南理解中类人认知水平的评估仍不充分。图文指南是一种结合文字、图像和符号的视觉形式,旨在通过结构化信息提升人类理解效率,天然体现人类感知与认知特征。本文提出OIG-Bench,一个覆盖多个领域的图文指南理解综合基准。为降低人工标注成本,我们设计半自动化标注流程,多个智能体协作生成初步图像描述,辅助人工构建图文对。基于OIG-Bench,我们对29个前沿MLLM进行了全面评估,包括开源与闭源模型。结果表明,Qwen2.5-VL-72B表现最佳,整体准确率达77%。然而所有模型在语义理解和逻辑推理方面均存在显著不足,显示当前模型仍难以准确解析复杂图文关系。此外,多智能体标注系统在图像描述任务上超越所有被测模型,展现出作为高质量图像描述生成器及未来数据集构建工具的潜力。数据集已公开于https://github.com/XiejcSYSU/OIG-Bench。

原文摘要 · Abstract (English)

Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities. However, evaluating their capacity for human-like understanding in One-Image Guides remains insufficiently explored. One-Image Guides are a visual format combining text, imagery, and symbols to present reorganized and structured information for easier comprehension, which are specifically designed for human viewing and inherently embody the characteristics of human perception and understanding. Here, we present OIG-Bench, a comprehensive benchmark focused on One-Image Guide understanding across diverse domains. To reduce the cost of manual annotation, we developed a semi-automated annotation pipeline in which multiple intelligent agents collaborate to generate preliminary image descriptions, assisting humans in constructing image-text pairs. With OIG-Bench, we have conducted a comprehensive evaluation of 29 state-of-the-art MLLMs, including both proprietary and open-source models. The results show that Qwen2.5-VL-72B performs the best among the evaluated models, with an overall accuracy of 77%. Nevertheless, all models exhibit notable weaknesses in semantic understanding and logical reasoning, indicating that current MLLMs still struggle to accurately interpret complex visual-text relationships. In addition, we also demonstrate that the proposed multi-agent annotation system outperforms all MLLMs in image captioning, highlighting its potential as both a high-quality image description generator and a valuable tool for future dataset construction. Datasets are available at https://github.com/XiejcSYSU/OIG-Bench.

多模态图文理解基准测试智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。