为测试测量领域量化大模型智能水平,建立可复现的评估基准。
TMIQ: Quantifying Test and Measurement Domain Intelligence in Large Language Models
- 构建TMIQ基准,覆盖电子工程任务的多维度评估
- 最佳模型SCPI命令匹配准确率73%,首项排名正确率33%
- 提供命令行工具支持定制化测试,适合工业级应用选型
测试测量领域对精度和效率要求严苛,正逐步引入生成式AI提升数据分析、自动化与决策能力。大型语言模型(LLMs)在测试自动化中展现巨大潜力,但其在该专业领域的评估仍不充分。为此,本文提出测试测量智商(TMIQ)基准,用于量化评估LLMs在电子工程任务中的表现。TMIQ涵盖SCPI指令匹配准确率、排序响应评估、思维链推理(CoT)及输出格式对性能的影响等多维指标。实验表明,各模型表现差异显著,精确匹配准确率介于56%至73%,最优模型首项匹配正确率约33%。同时评估了令牌使用量、成本效益与响应时间,揭示准确率与效率间的权衡。此外,本文还提供命令行接口(CLI)工具,支持按相同方法生成数据集,实现针对特定场景的模型评估。TMIQ与CLI工具为生产环境中的模型评估提供了严谨、可复现的方案,有助于持续监控与优化,推动测试测量领域模型选型创新。
原文摘要 · Abstract (English)
The Test and Measurement domain, known for its strict requirements for accuracy and efficiency, is increasingly adopting Generative AI technologies to enhance the performance of data analysis, automation, and decision-making processes. Among these, Large Language Models (LLMs) show significant promise for advancing automation and precision in testing. However, the evaluation of LLMs in this specialized area remains insufficiently explored. To address this gap, we introduce the Test and Measurement Intelligence Quotient (TMIQ), a benchmark designed to quantitatively assess LLMs across a wide range of electronic engineering tasks. TMIQ offers a comprehensive set of scenarios and metrics for detailed evaluation, including SCPI command matching accuracy, ranked response evaluation, Chain-of-Thought Reasoning (CoT), and the impact of output formatting variations required by LLMs on performance. In testing various LLMs, our findings indicate varying levels of proficiency, with exact SCPI command match accuracy ranging from around 56% to 73%, and ranked matching first-position scores achieving around 33% for the best-performing model. We also assess token usage, cost-efficiency, and response times, identifying trade-offs between accuracy and operational efficiency. Additionally, we present a command-line interface (CLI) tool that enables users to generate datasets using the same methodology, allowing for tailored assessments of LLMs. TMIQ and the CLI tool provide a rigorous, reproducible means of evaluating LLMs for production environments, facilitating continuous monitoring and identifying strengths and areas for improvement, and driving innovation in their selections for applications within the Test and Measurement industry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。