arXiv:2606.01820cs.CL2026-06

用大模型自动标注口语中的语法错误,省时省力。

TalkTag: Fine-Grained Morphosyntactic Error Annotation for Transcribed Speech

  • 基于大模型微调,自动分析儿童口语中的细粒度语法错误
  • 在数据稀缺下仍保持高精度,识别出复杂歧义情况
  • 适合语言研究、临床评估等需要大量标注的场景

细粒度形态句法错误标注在临床和发育语言研究中至关重要,但传统方法耗时耗力、依赖专家且难以扩展。本文提出 TalkTag,一个基于大模型的轻量级工具,经过微调后可自动化处理口语转录文本中的 CHAT 风格错误标注。该系统在极低资源条件下(使用儿童叙事数据)训练,验证了在低资源环境下进行语言分析的可行性。评估显示,TalkTag 能生成精准的标注结果,并有效识别出因语言歧义导致的自动化标注难题。总体而言,TalkTag 为手动标注提供了可扩展的替代方案,为形态句法错误标注提供了切实可行的技术支持。

原文摘要 · Abstract (English)

Fine-grained morphosyntactic error annotation is important in clinical and developmental language research, yet it is labour-intensive, expert-dependent, and difficult to scale. We present TalkTag, an LLM-based lightweight tool fine-tuned to automate CHAT-style error annotation in spoken-language transcripts. Developed under conditions of extreme data scarcity using children's narrative data, the system shows the feasibility of linguistic analysis in low-resource settings. Our evaluation demonstrates that TalkTag produces encouragingly precise annotation while effectively identifying instances where linguistic ambiguity makes automated tagging genuinely complex. In summary, with TalkTag, we provide a scalable alternative to manual error annotation and practically viable support for morphosyntactic error annotation.

语音分析大模型应用语法错误检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。