评测大模型处理多模态多结构数据的代码能力,发现顶尖模型仍有巨大提升空间。
BabelBench: An Omni Benchmark for Code-Driven Analysis of Multimodal and Multistructured Data
- 构建包含247个任务的统一评估框架,要求模型结合代码执行进行多类型推理。
- 测试显示即使GPT-4在逻辑推理和调试上仍存在明显短板。
- 适合关注大模型复杂任务泛化与系统性推理能力的研究者使用。
大型语言模型(LLMs)在多个领域日益关键,尤其在处理复杂数据类型方面。这包括结构化数据处理(如ChartQA和ChatGPT-Ada)以及多模态非结构化数据处理(如视觉问答VQA)。尽管如此,这些多样化数据处理场景仍缺乏统一的评估方法。为此,我们提出BabelBench,一个创新的基准框架,用于评估大模型在多模态多结构数据中通过代码执行进行分析的能力。该框架包含247个精心设计的问题,挑战模型在感知、常识推理、逻辑推理等方面的能力。除基本的多模态理解、结构化数据处理和代码生成外,这些任务还要求具备探索、规划、推理和调试等高级能力。我们在BabelBench上的实验表明,即使是最先进的模型如ChatGPT-4,也仍有显著改进空间。由此获得的洞察为社区未来研究提供了重要指引。基准数据可于https://github.com/FFD8FFE/babelbench获取。
原文摘要 · Abstract (English)
Large language models (LLMs) have become increasingly pivotal across various domains, especially in handling complex data types. This includes structured data processing, as exemplified by ChartQA and ChatGPT-Ada, and multimodal unstructured data processing as seen in Visual Question Answering (VQA). These areas have attracted significant attention from both industry and academia. Despite this, there remains a lack of unified evaluation methodologies for these diverse data handling scenarios. In response, we introduce BabelBench, an innovative benchmark framework that evaluates the proficiency of LLMs in managing multimodal multistructured data with code execution. BabelBench incorporates a dataset comprising 247 meticulously curated problems that challenge the models with tasks in perception, commonsense reasoning, logical reasoning, and so on. Besides the basic capabilities of multimodal understanding, structured data processing as well as code generation, these tasks demand advanced capabilities in exploration, planning, reasoning and debugging. Our experimental findings on BabelBench indicate that even cutting-edge models like ChatGPT 4 exhibit substantial room for improvement. The insights derived from our comprehensive analysis offer valuable guidance for future research within the community. The benchmark data can be found at https://github.com/FFD8FFE/babelbench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。