梳理传统方法在自然语言处理中的现状与应用价值。
A Structured Literature Review on Traditional Approaches in Current Natural Language Processing
- 从五个任务场景出发,系统评估传统方法的使用情况。
- 发现传统模型仍广泛用于流水线、基线或主模型中。
- 为研究者提供传统方法适用场景的实用参考。
近年来神经网络和大语言模型的兴起重塑了自然语言处理领域,推动了典型语言任务的新方法并取得主流成功。尽管大模型表现优异,但仍存在诸多不足。本文聚焦五个应用场景——应用情景分类、信息与关系抽取、文本简化及文本摘要,评估当前技术前沿,并探讨传统方法的未来潜力与合理应用场景。在明确定义传统技术特征后,调查其在近期论文中的实际使用情况。结果表明,这五个场景中均可见传统模型的身影:或作为处理流程的一部分,或作为核心模型的对比基准,或作为论文的主要模型。完整统计数据见 https://zenodo.org/records/13683801。
原文摘要 · Abstract (English)
The continued rise of neural networks and large language models in the more recent past has altered the natural language processing landscape, enabling new approaches towards typical language tasks and achieving mainstream success. Despite the huge success of large language models, many disadvantages still remain and through this work we assess the state of the art in five application scenarios with a particular focus on the future perspectives and sensible application scenarios of traditional and older approaches and techniques. In this paper we survey recent publications in the application scenarios classification, information and relation extraction, text simplification as well as text summarization. After defining our terminology, i.e., which features are characteristic for traditional techniques in our interpretation for the five scenarios, we survey if such traditional approaches are still being used, and if so, in what way they are used. It turns out that all five application scenarios still exhibit traditional models in one way or another, as part of a processing pipeline, as a comparison/baseline to the core model of the respective paper, or as the main model(s) of the paper. For the complete statistics, see https://zenodo.org/records/13683801
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。