用多维标签体系升级土耳其语学习者语料库,让错误分析更精细。
From Labels to Facets: Building a Taxonomically Enriched Turkish Learner Corpus
- 基于新提出的分面分类法,实现半自动化多维度标注
- 分面准确率达95.86%,显著提升语料信息丰富度
- 适合语言学研究者和二语习得领域学者使用
现有学习者语料库多采用单一扁平标签体系,难以分离多种语言维度,限制了深层语言分析。本文提出一种基于新分面分类法的半自动化标注方法,并构建针对土耳其语的标注扩展框架。该框架能自动从已有扁平标签中推断出更多语言与元数据信息作为分面,实现更丰富的学习者上下文标注。系统评估显示分面级准确率达95.86%。构建的分面增强型语料库支持复杂查询与细粒度探索性分析,使研究者可从语言与教学双重维度深入探究学习者错误模式。本研究首次发布协作标注、分面增强的土耳其语学习者语料库,包含人工标注指南、优化标签集及标注扩展工具,为后续类似语料库的拓展提供范式。
原文摘要 · Abstract (English)
In terms of annotation structure, most learner corpora rely on holistic flat label inventories which, even when extensive, do not explicitly separate multiple linguistic dimensions. This makes linguistically deep annotation difficult and complicates fine-grained analyses aimed at understanding why and how learners produce specific errors. To address these limitations, this paper presents a semi-automated annotation methodology for learner corpora, built upon a recently proposed faceted taxonomy, and implemented through a novel annotation extension framework. The taxonomy provides a theoretically grounded, multi-dimensional categorization that captures the linguistic properties underlying each error instance, thereby enabling standardized, fine-grained, and interpretable enrichment beyond flat annotations. The annotation extension tool, implemented based on the proposed extension framework for Turkish, automatically extends existing flat annotations by inferring additional linguistic and metadata information as facets within the taxonomy to provide richer learner-specific context. It was systematically evaluated and yielded promising performance results, achieving a facet-level accuracy of 95.86%. The resulting taxonomically enriched corpus offers enhanced querying capabilities and supports detailed exploratory analyses across learner corpora, enabling researchers to investigate error patterns through complex linguistic and pedagogical dimensions. This work introduces the first collaboratively annotated and taxonomically enriched Turkish Learner Corpus, a manual annotation guideline with a refined tagset, and an annotation extender. As the first corpus designed in accordance with the recently introduced taxonomy, we expect our study to pave the way for subsequent enrichment efforts of existing error-annotated learner corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。