用双解析器一致性辅助韩语二语标注,提升效率与准确性。
Parser agreement and disagreement in L2 Korean UD: Implications for human-in-the-loop annotation

- 通过两个领域适配解析器的共识生成初步标注
- 解析器一致率与人工判断高度吻合,验证可行性
- 分歧集中在句法关系和分句边界等可优化领域
我们提出一种简化的人机协同工作流,用于第二语言(L2)韩语形态句法标注,利用两个领域适配解析器的一致性。首先评估解析器一致性是否可作为标注正确性的代理指标,将其与独立的人工判断进行对比。结果显示解析器与人工判断间存在强对应关系,支持半自动L2韩语UD标注的可行性。进一步分析表明,解析器分歧在语法关系区分和分句边界模糊等语言学可预测领域集中出现。尽管多数分歧可通过迭代模型优化解决,但部分分歧反映了处理L2韩语文本时解析与标注固有的深层表征挑战。
原文摘要 · Abstract (English)
We propose a simplified human-in-the-loop workflow for second language (L2) Korean morphosyntactic annotation by leveraging agreement between two domain-adapted parsers. We first evaluate whether parser agreement can serve as a proxy for annotation correctness by comparing it with independent human judgments. The results show strong correspondence between parser and human judgments, supporting the feasibility of semi-automatic L2-Korean UD annotation. Further analysis demonstrates that parser disagreements cluster in linguistically predictable domains such as grammatical-relation distinctions and clause-boundary ambiguity. While many disagreement cases are tractable for iterative model refinement, others reflect deeper representational challenges inherent in parsing and tagging L2-Korean corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。