arXiv:2412.00102cs.CVcs.CL2024-12被引 5

首个面向电子电路的多模态问答基准,评测大模型理解能力

ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?

  • 构建专用于数字电路的视觉问答数据集,覆盖626个题目
  • 发现当前多模态大模型在电子电路问题上表现有限
  • 适合教育科技与工程类AI研究者参考

多模态大语言模型(MLLMs)在处理多模态数据方面备受关注,能够增强对复杂问题的上下文理解。尽管其在视觉问答(VQA)等任务中表现出色,但在基础工程问题上仍存在困难,且缺乏针对数字电子学的专用训练数据集。为此,本文提出一个名为ElectroVizQA的基准数据集,专门用于评估MLLMs在本科数字电子课程常见电路问题上的表现。该数据集是首个面向数字电子学VQA任务的专用数据集,包含约626个视觉问题,全面覆盖数字电子学核心知识点。本研究系统评估了MLLMs在理解和解决数字电路问题方面的能力,揭示其在该专业领域中的优势与局限。通过引入此基准,旨在推动相关研究进展,促进MLLMs在工程教育中的应用,缩小性能差距,提升模型在技术领域的实用性。

原文摘要 · Abstract (English)

Multi-modal Large Language Models (MLLMs) are gaining significant attention for their ability to process multi-modal data, providing enhanced contextual understanding of complex problems. MLLMs have demonstrated exceptional capabilities in tasks such as Visual Question Answering (VQA); however, they often struggle with fundamental engineering problems, and there is a scarcity of specialized datasets for training on topics like digital electronics. To address this gap, we propose a benchmark dataset called ElectroVizQA specifically designed to evaluate MLLMs' performance on digital electronic circuit problems commonly found in undergraduate curricula. This dataset, the first of its kind tailored for the VQA task in digital electronics, comprises approximately 626 visual questions, offering a comprehensive overview of digital electronics topics. This paper rigorously assesses the extent to which MLLMs can understand and solve digital electronic circuit questions, providing insights into their capabilities and limitations within this specialized domain. By introducing this benchmark dataset, we aim to motivate further research and development in the application of MLLMs to engineering education, ultimately bridging the performance gap and enhancing the efficacy of these models in technical fields.

多模态电子电路视觉问答教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。