arXiv:2605.13663cs.CLcs.CY2026-05ACL

提出分层提示法,提升社交媒体假新闻分类在不同标注标准下的鲁棒性。

Fine-tuning with Hierarchical Prompting for Robust Propaganda Classification Across Annotation Schemas

论文配图:Fine-tuning with Hierarchical Prompting for Robust Propaganda Classification Across Annotation Schemas
图 1 · 摘自论文原文
  • 设计分层提示策略,先预测细粒度传播技巧再聚合结果。
  • 微调后模型性能显著提升,Qwen系列在多数据集上表现最优。
  • 新标注数据集揭示传播策略目标,适合真实场景的鲁棒检测研究。

社交媒体中的宣传内容检测面临文本短小、噪声大及标注一致性低等挑战。本文提出一种以意图为核心的新型传播技巧分类体系,并与已有高一致性的标准体系进行对比。基于四个语言模型(GPT-4.1-nano、Phi-4 14B、Qwen2.5-14B、Qwen3-14B),从模型组合、标注体系影响和提示策略三个维度评估分类性能。结果显示,微调对提升零样本基线至关重要,能揭示基础模型隐藏的方法差异。在不同标注体系下,Qwen系列整体表现最佳,Phi-4 14B始终优于GPT-4.1-nano。提出的分层提示方法(HiPP)在微调后尤其有效,对模糊且低一致性的新标签体系帮助更大,同时在简单体系上仍具竞争力。新构建的HQP数据集采用意图导向标注,为理解宣传战略目标提供更丰富视角,并为未来鲁棒检测研究提供挑战性基准。

原文摘要 · Abstract (English)

Propaganda detection in social media is challenging due to noisy, short texts and low annotation agreements. We introduce a new intent-focused taxonomy of propaganda techniques and compare it against an established, higher-agreement schema. Along three dimensions (model portfolio, schema effects, and prompting strategy) we evaluate the taxonomies as a classification task with the help of four language models (GPT-4.1-nano, Phi-4 14B, Qwen2.5-14B, Qwen3-14B). Our results show that fine-tuning is essential, since it transforms weak zero-shot baselines into competitive systems and reveals methodological differences that are hidden using base models. Across schemas, the Qwen models achieve the strongest overall performance, and Phi-4 14B consistently outperforms GPT-4.1-nano. Our hierarchical prompting method (HiPP), which predicts fine-grained techniques before aggregating them, is especially beneficial after fine-tuning and on the more ambiguous, low-agreement taxonomy, while remaining competitive on the simpler schema. The HQP dataset, annotated with the new intent-based labels, provides a richer lens on propaganda's strategic goals and a challenging benchmark for future work on robust, real-world detection.

传播识别提示工程微调数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。