arXiv:2607.28119cs.CLcs.SI2026-07

LLM在主观语言标注中表现优于训练中语言学家,可辅助数字人文研究。

Challenges in annotations by humans and LLMs: A case study of evaluative language

  • 用三种提示词比较LLM对评价性语言的分类能力,优化后达最佳效果
  • 微调后模型F1得分0.77,优于专业语言学家标注结果
  • 适合从事数字人文、语料标注与复杂理论分析的研究者参考

本文比较了语言学训练中的学生、受训语言学家以及大语言模型(LLMs)在标注评价性语言时的表现,聚焦于科学类口语语料(如英文TED演讲文本)中的评价理论(Appraisal theory)及其态度子系统(包括情感、判断、欣赏三类)。该任务具有高度主观性,是复杂标注的典型代表。首先评估人类在特定科学领域句子级标注表现;随后设计三个提示词,对比模型在自动分类中的性能。采用最优提示词评估三款LLM并进行微调,最终实现F1分数0.77。结果显示,模型表现优于受训语言学家,而训练中的语言学家则未达到高一致性。结论表明,LLM可有效辅助复杂标注任务,为数字人文研究中的理论标注与分析开辟新路径。

原文摘要 · Abstract (English)

In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.

语言标注评价理论LLM应用数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。