用大模型自动标注汉语叙事的结构层次,效率提升65%且可靠接近人工。
LLMs for automatic annotation of Mandarin narrative transcripts
- 基于提示工程让大模型识别汉语叙事的语篇结构层次
- 最佳模型与人工标注一致性达k=0.794,时间节省65%
- 适合需高效处理中文口语语料的研究者使用
语音转写文本的语言学标注对语言习得、语言障碍和社会语言学研究至关重要,但长期依赖人力且耗时。尽管大语言模型在自动化标注方面展现潜力,其在非英语语言中处理复杂语篇层面标注的能力仍缺乏研究。本研究以多语言叙事评估工具(MAIN)为基准,评估大模型在汉语口语叙事中识别故事语法元素层级结构(即叙事宏观结构)的能力。对比四款大模型与受训人工标注者,在儿童、青年及老年群体产生的叙事上进行测试。最优模型与人工标注者达成一致率k=0.794,接近人工间可靠性k=0.872,同时减少65%标注时间;而本地部署的轻量模型表现明显较差。标注难度随结构类型系统性变化,需精细语义区分的类别持续构成挑战。此外,青年群体叙事因词汇多样性高、语义模糊及单句内多重元素整合,导致模型可靠性下降。结果表明,大模型可有效支持非英语口语语料的语篇级标注,但仍需人类监督处理语义复杂的任务。本文提出的提示模板已开源供后续使用。
原文摘要 · Abstract (English)
Linguistic annotation of transcribed speech is essential for research in language acquisition, language disorders, and sociolinguistics, yet remains labor-intensive and time-consuming. While Large Language Models (LLMs) have shown promise in automating annotation tasks, their ability to handle complex discourse-level annotation in non-English languages remains understudied. This study evaluates whether LLMs can reliably annotate narrative macrostructure-the hierarchical organization of story grammar elements-in spoken Mandarin, using the Multilingual Assessment Instrument for Narratives (MAIN) as a testbed. We compared four LLMs against trained human annotators on narratives produced by children, young adults, and older adults. The best-performing model achieved agreement with human raters (k=.794) approaching human-human reliability levels (k=.872) while reducing annotation time by 65%, whereas the locally deployable lightweight model performed substantially worse. Annotation difficulty varied systematically by macrostructure element type, with categories requiring subtle semantic differentiation posing persistent challenges. Furthermore, model reliability decreased on young adult narratives, which exhibited greater lexical variation, semantic ambiguity, and multi-element integration within single utterances. These findings suggest that LLMs can effectively support discourse-level annotation in non-English spoken corpora, while highlighting the continued need for human oversight in semantically complex tasks. Our prompt templates are open sourced for future use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。