构建电子工程多模态评测集,揭示大模型在真实工程任务中的能力短板。
EEE-Bench: A Comprehensive Multimodal Electrical And Electronics Engineering Benchmark
- 设计2860个电子工程问题,覆盖10个子领域,融合复杂图文信息
- 模型平均准确率仅19.48%至46.78%,暴露严重能力不足
- 发现模型'视觉懒惰'现象:忽视图像直接依赖文本推理
大型语言模型(LLMs)和多模态模型(LMMs)在科学与数学等领域展现潜力,但在更具挑战性的工程场景中仍缺乏系统评估。为此,我们提出EEE-Bench,一个以电子电气工程(EEE)为测试场景的多模态基准,旨在评估LMMs解决实际工程任务的能力。该基准包含2860个精心设计的问题,覆盖模拟电路、控制系统等10个核心子领域。相较于其他领域基准,工程问题具有更强的视觉复杂性与解法非确定性特征,需模型深度融合图文信息以理解抽象电路图与系统图。我们对17种主流开源与闭源的LLMs和LMMs进行了全面量化评估与细粒度分析,结果显示当前基础模型在EEE任务中表现不佳,平均准确率介于19.48%至46.78%之间。进一步发现,模型普遍存在‘懒惰’现象:在技术图像推理中倾向于依赖文本忽略视觉上下文。综上,EEE-Bench不仅揭示了当前LMMs的关键缺陷,也为推动其在真实工程应用中的研究与改进提供了重要资源。
原文摘要 · Abstract (English)
Recent studies on large language models (LLMs) and large multimodal models (LMMs) have demonstrated promising skills in various domains including science and mathematics. However, their capability in more challenging and real-world related scenarios like engineering has not been systematically studied. To bridge this gap, we propose EEE-Bench, a multimodal benchmark aimed at assessing LMMs' capabilities in solving practical engineering tasks, using electrical and electronics engineering (EEE) as the testbed. Our benchmark consists of 2860 carefully curated problems spanning 10 essential subdomains such as analog circuits, control systems, etc. Compared to benchmarks in other domains, engineering problems are intrinsically 1) more visually complex and versatile and 2) less deterministic in solutions. Successful solutions to these problems often demand more-than-usual rigorous integration of visual and textual information as models need to understand intricate images like abstract circuits and system diagrams while taking professional instructions, making them excellent candidates for LMM evaluations. Alongside EEE-Bench, we provide extensive quantitative evaluations and fine-grained analysis of 17 widely-used open and closed-sourced LLMs and LMMs. Our results demonstrate notable deficiencies of current foundation models in EEE, with an average performance ranging from 19.48% to 46.78%. Finally, we reveal and explore a critical shortcoming in LMMs which we term laziness: the tendency to take shortcuts by relying on the text while overlooking the visual context when reasoning for technical image problems. In summary, we believe EEE-Bench not only reveals some noteworthy limitations of LMMs but also provides a valuable resource for advancing research on their application in practical engineering tasks, driving future improvements in their capability to handle complex, real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。