用大模型零样本标注社交媒体健康数据,效率更高但复杂任务易出错。
Scaling Public Health Text Annotation: Zero-Shot Learning vs. Crowdsourcing for Improved Efficiency and Labeling Accuracy
- 用大语言模型零样本标注微博内容,替代人工标注提升效率。
- 简单分类任务中模型表现接近专家,复杂任务准确率下降。
- 适合需要快速批量标注的公共卫生研究,不适用于深度专业判断。
公共健康研究人员越来越关注利用社交媒体数据研究健康相关行为,但手动标注这些数据既费时又昂贵。本研究探讨了使用大语言模型(LLMs)进行零样本标注,是否能匹配或超越传统众包标注在睡眠障碍、体力活动和久坐行为相关推文上的表现。设计了多种标注流程,比较领域专家、众包工作者与基于LLM的方法在不同提示工程策略下的标注结果。结果显示,LLMs在简单分类任务中可媲美人类表现,并显著缩短标注时间,但在需要更细致领域知识的任务中准确性下降。研究明确了自动化扩展与人类专长之间的权衡,揭示了在不损害标签质量的前提下,如何高效将基于LLM的标注整合进公共健康研究。
原文摘要 · Abstract (English)
Public health researchers are increasingly interested in using social media data to study health-related behaviors, but manually labeling this data can be labor-intensive and costly. This study explores whether zero-shot labeling using large language models (LLMs) can match or surpass conventional crowd-sourced annotation for Twitter posts related to sleep disorders, physical activity, and sedentary behavior. Multiple annotation pipelines were designed to compare labels produced by domain experts, crowd workers, and LLM-driven approaches under varied prompt-engineering strategies. Our findings indicate that LLMs can rival human performance in straightforward classification tasks and significantly reduce labeling time, yet their accuracy diminishes for tasks requiring more nuanced domain knowledge. These results clarify the trade-offs between automated scalability and human expertise, demonstrating conditions under which LLM-based labeling can be efficiently integrated into public health research without undermining label quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。