构建克罗地亚新闻标题点击诱饵数据集,对比微调与提示学习效果。
What Makes You CLIC: Detection of Croatian Clickbait Headlines
- 基于BERTić模型在自建数据集上微调检测点击诱饵。
- 近半数标题为点击诱饵,微调模型优于通用大模型。
- 适合关注低资源语言信息质量与虚假标题识别的研究者。
在线新闻媒体主要依赖广告收入,促使记者撰写夸张、吸引眼球的标题,即点击诱饵。自动识别点击诱饵对维护数字媒体信息质量和读者信任至关重要,需结合上下文理解与世界知识。针对低资源语言,当前尚不明确微调方法与上下文学习(ICL)哪种更优。本文构建了CLIC数据集,涵盖克罗地亚20年间的主流与边缘新闻机构标题。我们对BERTić模型进行微调,并与使用克罗地亚语和英语提示的LLM-based ICL方法进行比较。结果表明,近一半分析标题属于点击诱饵,且微调模型性能优于通用大模型。同时,我们分析了点击诱饵的语言特征。
原文摘要 · Abstract (English)
Online news outlets operate predominantly on an advertising-based revenue model, compelling journalists to create headlines that are often scandalous, intriguing, and provocative -- commonly referred to as clickbait. Automatic detection of clickbait headlines is essential for preserving information quality and reader trust in digital media and requires both contextual understanding and world knowledge. For this task, particularly in less-resourced languages, it remains unclear whether fine-tuned methods or in-context learning (ICL) yield better results. In this paper, we compile CLIC, a novel dataset for clickbait detection of Croatian news headlines spanning a 20-year period and encompassing mainstream and fringe outlets. We fine-tune the BERTić model on this task and compare its performance to LLM-based ICL methods with prompts both in Croatian and English. Finally, we analyze the linguistic properties of clickbait. We find that nearly half of the analyzed headlines contain clickbait, and that finetuned models deliver better results than general LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。