提出轻量级文本归一化与语义解析框架,支持小数据场景本地快速部署。
Digestion Algorithm in Hierarchical Symbolic Forests: A Fast Text Normalization Algorithm and Semantic Parsing Framework for Specific Scenarios and Lightweight Deployment
- 基于分层符号森林的消化算法,融合组合数学与人类思维模式
- 模型规模与内存占用降低两个数量级,本地部署响应更快
- 适用于数据少、对可解释性要求高的风险敏感场景
文本归一化与语义解析在自然语言编程、改写、数据增强、专家系统构建、文本匹配等领域应用广泛。尽管大语言模型在深度学习中取得显著进展,但神经网络架构的可解释性差,影响其可信度,限制了在高风险场景中的部署。在特定领域数据稀缺时,快速获取大量标注数据困难,人工标注工作量巨大,且神经网络存在灾难性遗忘,导致数据利用率低。在需快速响应的场景中,模型密度高,本地部署困难,响应时间长。受乘法原理与人类思维模式启发,本文提出多层框架及算法——分层符号森林中的消化算法(DAHSF),整合文本归一化与语义解析流程。中文脚本语言“火兔智能开发平台V2.0”是该技术的重要测试与应用场景。DAHSF可在小数据场景下本地运行,模型大小与内存使用优化至少两个数量级,显著提升执行效率,具备良好优化前景。
原文摘要 · Abstract (English)
Text Normalization and Semantic Parsing have numerous applications in natural language processing, such as natural language programming, paraphrasing, data augmentation, constructing expert systems, text matching, and more. Despite the prominent achievements of deep learning in Large Language Models (LLMs), the interpretability of neural network architectures is still poor, which affects their credibility and hence limits the deployments of risk-sensitive scenarios. In certain scenario-specific domains with scarce data, rapidly obtaining a large number of supervised learning labels is challenging, and the workload of manually labeling data would be enormous. Catastrophic forgetting in neural networks further leads to low data utilization rates. In situations where swift responses are vital, the density of the model makes local deployment difficult and the response time long, which is not conducive to local applications of these fields. Inspired by the multiplication rule, a principle of combinatorial mathematics, and human thinking patterns, a multilayer framework along with its algorithm, the Digestion Algorithm in Hierarchical Symbolic Forests (DAHSF), is proposed to address these above issues, combining text normalization and semantic parsing workflows. The Chinese Scripting Language "Fire Bunny Intelligent Development Platform V2.0" is an important test and application of the technology discussed in this paper. DAHSF can run locally in scenario-specific domains on little datasets, with model size and memory usage optimized by at least two orders of magnitude, thus improving the execution speed, and possessing a promising optimization outlook.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。