用少样本提示让大模型检测文本偏见,多语言表现优异。
CEA-LIST at CheckThat! 2025: Evaluating LLMs as Detectors of Bias and Opinion in Text
- 通过精心设计的少样本提示,利用大模型实现多语言主观性检测。
- 在阿拉伯语、波兰语等多语言任务中获第一,对低质量数据更鲁棒。
- 无需微调,适合标注数据少或不一致的场景,实用性强。
本文提出一种基于大语言模型(LLMs)的少样本提示方法,用于多语言主观性检测。我们参加了CheckThat! 2025评测活动中的主题1任务。实验表明,当搭配精心设计的提示时,大模型在噪声大或质量差的数据环境下,性能可媲美甚至超越微调的小型语言模型(SLMs)。尽管尝试了辩论式提示、多种示例选择策略等高级提示工程技巧,但效果提升有限,标准少样本提示已足够有效。我们的系统在多个语言赛道中取得顶尖成绩,包括阿拉伯语和波兰语第一,意大利语、英语、德语及多语言赛道均进入前四。尤其在阿拉伯语数据集上表现突出,可能得益于对标注不一致性的强适应能力。结果表明,基于大模型的少样本学习在多语言情感分析任务中具有高效性与灵活性,是传统微调的有力替代方案,尤其适用于标注数据稀缺或不一致的情况。
原文摘要 · Abstract (English)
This paper presents a competitive approach to multilingual subjectivity detection using large language models (LLMs) with few-shot prompting. We participated in Task 1: Subjectivity of the CheckThat! 2025 evaluation campaign. We show that LLMs, when paired with carefully designed prompts, can match or outperform fine-tuned smaller language models (SLMs), particularly in noisy or low-quality data settings. Despite experimenting with advanced prompt engineering techniques, such as debating LLMs and various example selection strategies, we found limited benefit beyond well-crafted standard few-shot prompts. Our system achieved top rankings across multiple languages in the CheckThat! 2025 subjectivity detection task, including first place in Arabic and Polish, and top-four finishes in Italian, English, German, and multilingual tracks. Notably, our method proved especially robust on the Arabic dataset, likely due to its resilience to annotation inconsistencies. These findings highlight the effectiveness and adaptability of LLM-based few-shot learning for multilingual sentiment tasks, offering a strong alternative to traditional fine-tuning, particularly when labeled data is scarce or inconsistent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。