构建跨行业多模态评测基准,评估大模型在真实工业场景中的表现
MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark
- 人工构建21个工业领域数据集,每域50个问答对,确保专业性
- 涵盖需专业知识与非OCR任务,提升评测复杂度
- 支持中英文双语对比,适合工业应用研究者参考
随着多模态大语言模型(MLLMs)的快速发展,大量评测基准应运而生。然而,针对其在多样化工业应用场景下的综合评估仍显不足。本文提出MME-Industry,一个专为工业场景设计的新型多模态评测基准。该基准覆盖21个不同领域,包含1050个问答对(每领域50个),所有问答对均由领域专家手工制作并验证,确保数据完整性和安全性。通过引入无需OCR即可直接回答的问题,以及需要专业领域知识的任务,有效提升了评测难度。此外,基准提供中英文双语版本,支持跨语言能力对比分析。实验结果为MLLMs在实际工业应用中的表现提供了重要洞见,并指明了未来模型优化的研究方向。
原文摘要 · Abstract (English)
With the rapid advancement of Multimodal Large Language Models (MLLMs), numerous evaluation benchmarks have emerged. However, comprehensive assessments of their performance across diverse industrial applications remain limited. In this paper, we introduce MME-Industry, a novel benchmark designed specifically for evaluating MLLMs in industrial settings.The benchmark encompasses 21 distinct domain, comprising 1050 question-answer pairs with 50 questions per domain. To ensure data integrity and prevent potential leakage from public datasets, all question-answer pairs were manually crafted and validated by domain experts. Besides, the benchmark's complexity is effectively enhanced by incorporating non-OCR questions that can be answered directly, along with tasks requiring specialized domain knowledge. Moreover, we provide both Chinese and English versions of the benchmark, enabling comparative analysis of MLLMs' capabilities across these languages. Our findings contribute valuable insights into MLLMs' practical industrial applications and illuminate promising directions for future model optimization research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。