用大模型优化宫颈癌疫苗争议内容标注,发现提示学习更高效。
Optimizing Social Media Annotation of HPV Vaccine Skepticism and Misinformation Using Large Language Models: An Experimental Evaluation of In-Context Learning and Fine-Tuning Stance Detection Across Multiple Models
- 对比提示学习与微调,发现提示学习在识别疫苗态度上表现更好。
- 六组分层样本搭配详细提示可达到最佳检测效果。
- 不同大模型对提示条件敏感度不同,需针对性调整配置。
本文利用大语言模型(LLMs)实验评估在社交媒体上进行人乳头瘤病毒(HPV)疫苗相关言论立场检测的最优策略。研究比较了传统微调与新兴的上下文学习方法,系统性地改变提示工程策略,涵盖多种主流大模型及其变体(如GPT-4、Mistral、Llama3等),包括提示模板设计、样本采样方式和样本数量。结果表明:1)总体而言,上下文学习在识别HPV疫苗社交媒体立场方面优于微调;2)增加样本数量并不必然提升性能;3)不同大模型及其变体对上下文学习条件的敏感性存在差异。研究发现,针对HPV疫苗推文的最佳上下文学习配置为六组分层样本搭配详细上下文提示。本研究揭示了大模型在社交媒体立场与疑虑检测中的潜力,并提供了可操作的应用路径。
原文摘要 · Abstract (English)
This paper leverages large-language models (LLMs) to experimentally determine optimal strategies for scaling up social media content annotation for stance detection on HPV vaccine-related tweets. We examine both conventional fine-tuning and emergent in-context learning methods, systematically varying strategies of prompt engineering across widely used LLMs and their variants (e.g., GPT4, Mistral, and Llama3, etc.). Specifically, we varied prompt template design, shot sampling methods, and shot quantity to detect stance on HPV vaccination. Our findings reveal that 1) in general, in-context learning outperforms fine-tuning in stance detection for HPV vaccine social media content; 2) increasing shot quantity does not necessarily enhance performance across models; and 3) different LLMs and their variants present differing sensitivity to in-context learning conditions. We uncovered that the optimal in-context learning configuration for stance detection on HPV vaccine tweets involves six stratified shots paired with detailed contextual prompts. This study highlights the potential and provides an applicable approach for applying LLMs to research on social media stance and skepticism detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。