用大模型生成稳定认知扭曲标注,提升主观任务的训练效果。
Towards Consistent Detection of Cognitive Distortions: LLM-Based Annotation and Dataset-Agnostic Evaluation
- 多轮独立大模型标注揭示稳定标签模式,缓解主观性问题。
- GPT-4标注一致性达Fleiss's Kappa=0.78,优于人工标注。
- 提出基于Cohen's kappa的跨数据集评估框架,适合公平比较。
文本中的认知扭曲自动检测因主观性强而困难,即使专家人工标注也存在低一致性,导致标注不可靠。本文探索使用大语言模型(LLMs)作为一致且可靠的标注者,提出多次独立运行大模型可揭示尽管任务主观但仍稳定的标注模式。为公平比较不同数据集训练模型的表现,引入基于Cohen's kappa的无数据集依赖评估框架,克服传统F1分数在跨数据集比较中的局限。实验表明,GPT-4生成的标注具有一致性(Fleiss's Kappa = 0.78),基于这些标注训练的模型在测试集上表现优于基于人工标注训练的模型。结果表明,大模型可提供可扩展、内部一致的训练数据生成方式,支持主观自然语言处理任务的强下游性能。
原文摘要 · Abstract (English)
Text-based automated Cognitive Distortion detection is a challenging task due to its subjective nature, with low agreement scores observed even among expert human annotators, leading to unreliable annotations. We explore the use of Large Language Models (LLMs) as consistent and reliable annotators, and propose that multiple independent LLM runs can reveal stable labeling patterns despite the inherent subjectivity of the task. Furthermore, to fairly compare models trained on datasets with different characteristics, we introduce a dataset-agnostic evaluation framework using Cohen's kappa as an effect size measure. This methodology allows for fair cross-dataset and cross-study comparisons where traditional metrics like F1 score fall short. Our results show that GPT-4 can produce consistent annotations (Fleiss's Kappa = 0.78), resulting in improved test set performance for models trained on these annotations compared to those trained on human-labeled data. Our findings suggest that LLMs can offer a scalable and internally consistent alternative for generating training data that supports strong downstream performance in subjective NLP tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。