对比微调与上下文学习,发现微调在虚假内容检测中更有效。
Are LLMs Enough for Hyperpartisan, Fake, Polarized and Harmful Content Detection? Evaluating In-Context Learning vs. Fine-Tuning
- 比较微调与多种上下文学习策略,评估其在多语言任务中的表现。
- 微调模型在10个数据集上普遍优于上下文学习,即使使用大模型亦然。
- 适合关注社交媒体虚假信息检测的研究者与实践者参考。
在线平台虚假新闻、极化、政治偏见及有害内容的传播已成严重问题。尽管大语言模型展现出潜力,但尚无研究全面评估其在不同模型、使用方式和语言下的表现。本研究系统评估了大语言模型在检测极化新闻、虚假新闻及政治偏见方面的适应范式,覆盖10个数据集和5种语言(英语、西班牙语、葡萄牙语、阿拉伯语、保加利亚语),包含二分类与多分类场景。实验涵盖参数高效的微调、零样本提示、代码本、少量样本(随机与基于确定性点过程的多样化样本)以及思维链提示等多种策略。结果表明,在多数情况下,上下文学习性能低于微调。这一发现凸显了即使面对最大模型(如LlaMA3.1-8b-Instruct、Mistral-Nemo-Instruct-2407、Qwen2.5-7B-Instruct),在特定任务上进行微调仍具关键价值。
原文摘要 · Abstract (English)
The spread of fake news, polarizing, politically biased, and harmful content on online platforms has been a serious concern. With large language models becoming a promising approach, however, no study has properly benchmarked their performance across different models, usage methods, and languages. This study presents a comprehensive overview of different Large Language Models adaptation paradigms for the detection of hyperpartisan and fake news, harmful tweets, and political bias. Our experiments spanned 10 datasets and 5 different languages (English, Spanish, Portuguese, Arabic and Bulgarian), covering both binary and multiclass classification scenarios. We tested different strategies ranging from parameter efficient Fine-Tuning of language models to a variety of different In-Context Learning strategies and prompts. These included zero-shot prompts, codebooks, few-shot (with both randomly-selected and diversely-selected examples using Determinantal Point Processes), and Chain-of-Thought. We discovered that In-Context Learning often underperforms when compared to Fine-Tuning a model. This main finding highlights the importance of Fine-Tuning even smaller models on task-specific settings even when compared to the largest models evaluated in an In-Context Learning setup - in our case LlaMA3.1-8b-Instruct, Mistral-Nemo-Instruct-2407 and Qwen2.5-7B-Instruct.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。