标注新加坡口语英语的词性标签很困难,自动工具准确率仅80%。
Limpeh ga li gong: Challenges in Singlish Annotations
- 由母语者构建带中英文对照和词性标注的平行数据集
- 现有模型在人工标注上准确率仅约80%
- 适合研究方言化语言处理、多语种自然语言任务的研究者
新加坡口语英语(Singlish)是多元文化背景下形成的口头交流语言。本文开展基础自然语言处理任务:对Singlish句子进行词性标注。我们构建了一个平行数据集,包含直接英文翻译及母语者标注的词性标签。实验表明,基于转移和Transformer的自动标注器在与人工标注对比时,准确率仅为约80%,说明该语言的计算分析仍有提升空间。本文揭示了标注Singlish的主要挑战:形式与语义不一致、高度依赖上下文的语气词、独特的结构表达,以及不同媒介间的语言变异。本研究的任务定义、标注结果和实验数据反映了由多种方言融合而成的口语语言分析难点,为未来超越词性标注的研究奠定基础。
原文摘要 · Abstract (English)
Singlish, or Colloquial Singapore English, is a language formed from oral and social communication within multicultural Singapore. In this work, we work on a fundamental Natural Language Processing (NLP) task: Parts-Of-Speech (POS) tagging of Singlish sentences. For our analysis, we build a parallel Singlish dataset containing direct English translations and POS tags, with translation and POS annotation done by native Singlish speakers. Our experiments show that automatic transition- and transformer- based taggers perform with only $\sim 80\%$ accuracy when evaluated against human-annotated POS labels, suggesting that there is indeed room for improvement on computation analysis of the language. We provide an exposition of challenges in Singlish annotation: its inconsistencies in form and semantics, the highly context-dependent particles of the language, its structural unique expressions, and the variation of the language on different mediums. Our task definition, resultant labels and results reflects the challenges in analysing colloquial languages formulated from a variety of dialects, and paves the way for future studies beyond POS tagging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。