轻量多专家系统提升工程文本生成准确率,降低算力需求。
A Lightweight Multi-Expert Generative Language Model System for Engineering Information and Knowledge Extraction
- 构建图结构小模型专家集群,每个节点为特定领域微调的小语言模型。
- 精确匹配率是传统微调的3倍,微调速度提升1.7倍。
- 适合中小工程企业部署,支持分布式AI,减少对昂贵算力依赖。
尽管大语言模型领域适应技术取得进展,但现有方法仍计算开销大,且存在幻觉问题。多数方法未优先考虑微调与推理的资源消耗。虽然幻觉随新模型发布逐渐减少,但在工程场景中仍普遍存在,而生成结构清晰、错误极少的文本至关重要。本文提出轻量级适应方案Small Language Graph(SLG),其采用图结构,每个节点为针对特定简洁文本微调的小语言模型。实验表明,SLG在精确匹配指标上优于传统微调方法3倍,微调速度比独立大模型快1.7倍。该成果使中小型工程公司可无需投入昂贵算力即可安全使用生成式AI。图架构与小规模专家节点也为分布式人工智能系统提供可能,或可缓解全球对高成本集中式计算集群的依赖。
原文摘要 · Abstract (English)
Despite recent advancements in domain adaptation techniques for large language models, these methods remain computationally intensive, and the resulting models can still exhibit hallucination issues. Most existing adaptation methods do not prioritize reducing the computational resources required for fine-tuning and inference of language models. Hallucination issues have gradually decreased with each new model release. However, they remain prevalent in engineering contexts, where generating well-structured text with minimal errors and inconsistencies is critical. This work introduces a novel approach called the Small Language Graph (SLG), which is a lightweight adaptation solution designed to address the two key challenges outlined above. The system is structured in the form of a graph, where each node represents a lightweight expert - a small language model fine-tuned on specific and concise texts. The results of this study have shown that SLG was able to surpass conventional fine-tuning methods on the Exact Match metric by 3 times. Additionally, the fine-tuning process was 1.7 times faster compared to that of a larger stand-alone language model. These findings introduce a potential for small to medium-sized engineering companies to confidently use generative AI technologies, such as LLMs, without the necessity to invest in expensive computational resources. Also, the graph architecture and the small size of expert nodes offer a possible opportunity for distributed AI systems, thus potentially diverting the global need for expensive centralized compute clusters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。