arXiv:2506.03922cs.CLcs.AI2025-06被引 21

首个评估多模态大模型人文社科能力的跨语言基准

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

  • 构建多领域专家与自动化工具协作的数据生成流程
  • 包含1.3万+样本,覆盖6大人文社科类别与6种联合国语言
  • 揭示当前顶尖模型在跨学科思维上的明显短板

多模态大语言模型(MLLMs)在多个领域展现出巨大潜力,但现有评测基准主要聚焦于理工科常见的垂直逻辑推理和通用知识,忽视了人文学科与社会科学(HSS)的独特需求。HSS任务更依赖横向、跨领域的综合思考,以及抽象概念与视觉表征间的深层关联,这对MLLMs构成独特挑战。为此,我们提出HSSBench,一个专为评估MLLM在人文学科与社会科学任务中表现而设计的多语言基准,涵盖联合国六种官方语言。我们还开发了一套新型数据生成流水线,由多位领域专家与自动化代理协同生成并迭代优化每条样本。HSSBench包含超过13,000个精心设计的样本,覆盖六大核心类别。我们在该基准上对20余种主流MLLM进行了评测,结果表明,即便最先进的模型也面临显著挑战。我们希望该基准能推动学术界进一步提升MLLM的跨学科推理能力,尤其在知识融合与跨领域连接方面。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchmarks for evaluating MLLMs primarily emphasize general knowledge and vertical step-by-step reasoning typical of STEM disciplines, while overlooking the distinct needs and potential of the Humanities and Social Sciences (HSS). Tasks in the HSS domain require more horizontal, interdisciplinary thinking and a deep integration of knowledge across related fields, which presents unique challenges for MLLMs, particularly in linking abstract concepts with corresponding visual representations. Addressing this gap, we present HSSBench, a dedicated benchmark designed to assess the capabilities of MLLMs on HSS tasks in multiple languages, including the six official languages of the United Nations. We also introduce a novel data generation pipeline tailored for HSS scenarios, in which multiple domain experts and automated agents collaborate to generate and iteratively refine each sample. HSSBench contains over 13,000 meticulously designed samples, covering six key categories. We benchmark more than 20 mainstream MLLMs on HSSBench and demonstrate that it poses significant challenges even for state-of-the-art models. We hope that this benchmark will inspire further research into enhancing the cross-disciplinary reasoning abilities of MLLMs, especially their capacity to internalize and connect knowledge across fields.

多模态人文社科跨学科评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。