对比手工与工业生成语言资源的优劣,探索最佳实践路径。
Lexicons and grammars for language processing: industrial or handcrafted products?
- 对比手工构建与自动化生成语言资源的方法差异。
- 指出手工资源信息更丰富,但耗时长,自动化可提升效率。
- 适合语言学与计算机科学交叉研究者参考。
近年来,语言处理对语言数据的需求持续增长,这些数据现被称为语言资源。目前广泛使用的资源包括文本语料库(如Brown Corpus、Penn Treebank),以及近年发展的电子词典(WordNet、FrameNet、VerbNet、ComLex、Lexicon-Grammar等)和形式化语法(如TAG)。大多数词典与语法仍依赖人工构建,而语料库的构建则高度自动化。然而,越来越多语言处理专家认识到,词典与语法的信息含量高于语料库,因而能支持更复杂的处理任务。这种差异可能源于构建时间:语言学家手工打造的资源更具信息量。未来演化或朝两个方向发展:一是语言技术专家逐渐适应使用高信息量、复杂的手工资源;二是词典与语法的构建实现自动化与工业化,这是当前主流趋势。两种路径均已推进,且存在张力。语言学家与计算机科学家的关系取决于未来走向——前者需大量语言学人才,后者依赖工程师方案。本文分析典型语言资源实例,探讨手工、工业生成或二者结合哪种方式能带来最优结果或最具可行性。
原文摘要 · Abstract (English)
During the recent years, the use of linguistic data for language processing increased progressively. Such data are now commonly called language resources. Most of the language resources used for this purpose are collections of texts as the Brown Corpus and the Penn Treebank, but electronic lexicons (WordNet, FrameNet, VerbNet, ComLex, Lexicon-Grammar...) and formal grammars (TAG...) developed recently. Most processes of construction of lexicons and grammars are manual, whereas the construction of corpora has always been highly automated. However, more and more specialists of language processing realize that the information content of lexicons and grammars is richer than that of corpora, and hence the former make more elaborate processing possible. The difference in construction time is likely to be connected with the difference in information content: the handcrafting of lexicons and grammars by linguists would make them more informative than automatically generated data. This situation can evolve into two directions: either specialists of language technology get progressively used to handling manually constructed resources, which are more informative and more complex, or the process of construction of lexicons and grammars is automated and industrialized, which is the mainstream perspective. Both evolutions are already in progress, and a tension exists between them. The relation between linguists and computer scientists depends on the future of these evolutions, since the first implies training and hiring numerous linguists, whereas the other depends essentially on solutions elaborated by computer engineers. The aim of this article is to analyse practical examples of the language resources in question, and to discuss about which of the two trends, handcrafting or generating industrially, or a combination of both, can give the best results or is the most realistic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。