arXiv:2506.12958cs.LG2025-06被引 9

为多模态大模型创建分领域的评估基准,助力通用人工智能发展

Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

  • 构建七大关键学科的评估框架,覆盖多领域应用场景
  • 系统梳理各领域基准与研究现状,揭示模型能力与挑战
  • 提供可检索的领域化基准资源,适合研究人员参考

大型语言模型(LLMs)因其强大的推理与问题解决能力,正被广泛应用于多个学科领域。为衡量其效能,已有多种基准用于评估其推理、理解与解题能力。尽管已有若干综述讨论了LLM评估与基准,但领域特异性分析仍不充分。本文提出一个涵盖七个关键学科的分类体系,全面回顾各领域中的LLM基准与综述论文,突出模型在特定场景下的能力与应用挑战。最后,我们将这些基准按领域整理归类,形成可供研究人员查阅的实用资源,旨在推动向通用人工智能(AGI)迈进。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being deployed across disciplines due to their advanced reasoning and problem solving capabilities. To measure their effectiveness, various benchmarks have been developed that measure aspects of LLM reasoning, comprehension, and problem-solving. While several surveys address LLM evaluation and benchmarks, a domain-specific analysis remains underexplored in the literature. This paper introduces a taxonomy of seven key disciplines, encompassing various domains and application areas where LLMs are extensively utilized. Additionally, we provide a comprehensive review of LLM benchmarks and survey papers within each domain, highlighting the unique capabilities of LLMs and the challenges faced in their application. Finally, we compile and categorize these benchmarks by domain to create an accessible resource for researchers, aiming to pave the way for advancements toward artificial general intelligence (AGI)

多模态模型评估基准领域分析综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。