利用大模型迎合心理生成对立观点,提升假标题检测效果
Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning
- 让大模型先生成支持与反对的两套理由,形成对比
- 在三个数据集上检测准确率超越现有方法
- 适合关注虚假信息、内容安全的研究者和从业者
网络内容泛滥加剧了对假标题(clickbait)的担忧,这类标题通过夸大或误导吸引点击。尽管大语言模型(LLMs)在识别此类内容方面具有潜力,但其表现常受‘迎合倾向’(Sycophancy)影响——即更倾向于生成符合用户预期的推理,而非真实合理的判断。本文提出一种新思路:不消除迎合倾向,而是将其转化为优势,引导模型生成支持与反对同一标题的对比性推理。为此,我们设计了自更新反向立场推理生成框架(SORG),无需真实标签即可生成高质量的“同意”与“反对”推理对。进一步构建基于对抗推理的点击诱骗检测模型(ORCD),采用三个BERT编码器分别表示标题及对应推理,并通过软标签驱动的对比学习增强鲁棒性。在三个基准数据集上的实验表明,该方法在检测性能上持续优于仅用提示(prompting)的LLM、微调的小型语言模型以及当前最先进的检测基线。
原文摘要 · Abstract (English)
The widespread proliferation of online content has intensified concerns about clickbait, deceptive or exaggerated headlines designed to attract attention. While Large Language Models (LLMs) offer a promising avenue for addressing this issue, their effectiveness is often hindered by Sycophancy, a tendency to produce reasoning that matches users' beliefs over truthful ones, which deviates from instruction-following principles. Rather than treating sycophancy as a flaw to be eliminated, this work proposes a novel approach that initially harnesses this behavior to generate contrastive reasoning from opposing perspectives. Specifically, we design a Self-renewal Opposing-stance Reasoning Generation (SORG) framework that prompts LLMs to produce high-quality agree and disagree reasoning pairs for a given news title without requiring ground-truth labels. To utilize the generated reasoning, we develop a local Opposing Reasoning-based Clickbait Detection (ORCD) model that integrates three BERT encoders to represent the title and its associated reasoning. The model leverages contrastive learning, guided by soft labels derived from LLM-generated credibility scores, to enhance detection robustness. Experimental evaluations on three benchmark datasets demonstrate that our method consistently outperforms LLM prompting, fine-tuned smaller language models, and state-of-the-art clickbait detection baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。