构建兼顾历史与现代语法研究的德语树库,创新文本分类与标注方法。
GiesKaNe: Bridging Past and Present in Grammatical Theory and Practical Application
- 融合人工与机器协作,实现深层句法标注与文本分类。
- 提出新型口语-书面语连续体分类法,提升文本选择科学性。
- 仅用通用工具(如表格)即可完成复杂标注流程,适合资源有限团队。
本文探讨吉森大学与卡塞尔大学联合开展的GiesKaNe项目在语料库建设中的需求。该项目是参考语料库、历史语料库及深层句法标注的树库,旨在连接历史与当代语料,保持跨时空语言研究的相关性。其编纂过程平衡创新与标准,兼顾项目内部目标与学术共同体利益。方法上采用人机协同机制,涵盖分词、归一化、句子界定、标注、解析及标注者间一致性等基础环节,并讨论语法模型、标注框架与现有事实标准的比较。文中提出一种基于概念性口语与书面语连续体的机器辅助文本分类新方法,为文本选取提供新视角。同时,提出从既有标注中推导事实标准标注的方法,调和标准化与创新之间的矛盾。文章展示,即使像GiesKaNe这样雄心勃勃的项目,也可借助现有研究基础设施完成,无需专用标注工具,仅通过合理运用简单电子表格与现有系统即可实现高效工作流。
原文摘要 · Abstract (English)
This article explores the requirements for corpus compilation within the GiesKaNe project (University of Giessen and Kassel, Syntactic Basic Structures of New High German). The project is defined by three central characteristics: it is a reference corpus, a historical corpus, and a syntactically deeply annotated treebank. As a historical corpus, GiesKaNe aims to establish connections with both historical and contemporary corpora, ensuring its relevance across temporal and linguistic contexts. The compilation process strikes the balance between innovation and adherence to standards, addressing both internal project goals and the broader interests of the research community. The methodological complexity of such a project is managed through a complementary interplay of human expertise and machine-assisted processes. The article discusses foundational topics such as tokenization, normalization, sentence definition, tagging, parsing, and inter-annotator agreement, alongside advanced considerations. These include comparisons between grammatical models, annotation schemas, and established de facto annotation standards as well as the integration of human and machine collaboration. Notably, a novel method for machine-assisted classification of texts along the continuum of conceptual orality and literacy is proposed, offering new perspectives on text selection. Furthermore, the article introduces an approach to deriving de facto standard annotations from existing ones, mediating between standardization and innovation. In the course of describing the workflow the article demonstrates that even ambitious projects like GiesKaNe can be effectively implemented using existing research infrastructure, requiring no specialized annotation tools. Instead, it is shown that the workflow can be based on the strategic use of a simple spreadsheet and integrates the capabilities of the existing infrastructure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。