arXiv:2509.17421cs.CLcs.MM2025-09EMNLP

首个中文多图理解基准,专为真实场景设计。

RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios

  • 基于真实用户生成内容构建,覆盖多样场景与图像结构。
  • 21个模型测试显示闭源模型仍面临挑战,开源模型平均差71.8%。
  • 适合研究中文多图理解的学者与开发者使用。

尽管已有多种多模态多图评估数据集出现,但它们主要基于英文,尚无中文多图数据集。为此,我们推出了首个中文多模态多图数据集RealBench,包含9393个样本和69910张图像。RealBench通过引入真实用户生成内容,确保与实际应用的高度相关性,并涵盖丰富场景、图像分辨率与结构,进一步提升多图理解难度。我们使用21种不同规模的多模态大模型(包括支持多图输入的闭源模型及开源视觉与视频模型)对RealBench进行了全面评估。实验结果表明,即使是最强大的闭源模型在处理中文多图场景时仍面临挑战,且开源视觉/视频模型与闭源模型之间平均存在约71.8%的性能差距。这些结果说明RealBench为探索中文语境下的多图理解能力提供了重要研究基础。

原文摘要 · Abstract (English)

While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset. To fill this gap, we introduce RealBench, the first Chinese multimodal multi-image dataset, which contains 9393 samples and 69910 images. RealBench distinguishes itself by incorporating real user-generated content, ensuring high relevance to real-world applications. Additionally, the dataset covers a wide variety of scenes, image resolutions, and image structures, further increasing the difficulty of multi-image understanding. Ultimately, we conduct a comprehensive evaluation of RealBench using 21 multimodal LLMs of different sizes, including closed-source models that support multi-image inputs as well as open-source visual and video models. The experimental results indicate that even the most powerful closed-source models still face challenges when handling multi-image Chinese scenarios. Moreover, there remains a noticeable performance gap of around 71.8\% on average between open-source visual/video models and closed-source models. These results show that RealBench provides an important research foundation for further exploring multi-image understanding capabilities in the Chinese context.

多图理解中文数据集基准测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。