arXiv:2508.20420cs.CL2025-08被引 3

首个面向民航维修的工业级大模型评测基准,助力智能维修系统落地。

CAMB: A comprehensive industrial LLM benchmark on civil aviation maintenance

  • 构建专用于民航维修领域的工业级大模型评测基准
  • 揭示现有大模型在维修知识与复杂推理中的显著短板
  • 适合从事航空AI、RAG优化及领域模型开发的研究者

民航维修领域对行业标准要求严苛,其维护流程与故障排查属于高知识密度、强推理需求的任务。针对当前大语言模型(LLM)评估体系在该垂直领域缺乏专门工具的问题,本文提出并构建了一个工业级基准,旨在标准化评估LLM在民航维修场景下的能力,识别出其在领域知识与复杂推理方面的具体缺陷。通过定位这些不足,为领域微调、RAG优化或提示工程等改进策略提供依据,推动更智能的维修解决方案发展。本工作弥补了现有评估主要聚焦数学与编码推理的局限,同时鉴于检索增强生成(RAG)是当前主流应用方案,我们利用该基准评估了多个知名向量嵌入模型与LLM在实际维修场景中的表现。实验分析验证了基准的有效性,并开源了评测基准与代码:https://github.com/CamBenchmark/cambenchmark

原文摘要 · Abstract (English)

Civil aviation maintenance is a domain characterized by stringent industry standards. Within this field, maintenance procedures and troubleshooting represent critical, knowledge-intensive tasks that require sophisticated reasoning. To address the lack of specialized evaluation tools for large language models (LLMs) in this vertical, we propose and develop an industrial-grade benchmark specifically designed for civil aviation maintenance. This benchmark serves a dual purpose: It provides a standardized tool to measure LLM capabilities within civil aviation maintenance, identifying specific gaps in domain knowledge and complex reasoning. By pinpointing these deficiencies, the benchmark establishes a foundation for targeted improvement efforts (e.g., domain-specific fine-tuning, RAG optimization, or specialized prompt engineering), ultimately facilitating progress toward more intelligent solutions within civil aviation maintenance. Our work addresses a significant gap in the current LLM evaluation, which primarily focuses on mathematical and coding reasoning tasks. In addition, given that Retrieval-Augmented Generation (RAG) systems are currently the dominant solutions in practical applications , we leverage this benchmark to evaluate existing well-known vector embedding models and LLMs for civil aviation maintenance scenarios. Through experimental exploration and analysis, we demonstrate the effectiveness of our benchmark in assessing model performance within this domain, and we open-source this evaluation benchmark and code to foster further research and development:https://github.com/CamBenchmark/cambenchmark

大模型评测民航维修RAG工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。