arXiv:2506.09381cs.CL2025-06被引 3

用机器学习区分新闻标题质量高低,准确率达90.3%。

Binary classification for perceived quality of headlines and links on worldwide news websites, 2018-2024

  • 基于115个语言特征与1.2亿条新闻数据训练分类模型。
  • 微调DistilBERT达90.3%准确率,优于传统集成方法。
  • 适合关注信息可信度、内容审核的研究者与从业者。

在线新闻的泛滥可能导致低质量新闻标题/链接的广泛传播。为此,我们探究了能否自动区分感知质量较低与较高的新闻标题/链接。在2018-2024年间全球新闻网站的57,544,214条链接/标题组成的二元平衡数据集上(每类28,772,107条),提取了115个语言学特征,并基于专家共识评分对每条文本打标。传统集成方法中,袋装分类器表现最佳(88.1%准确率,88.3% F1,80/20训练/测试分割)。微调后的DistilBERT达到最高准确率(90.3%,80/20分割),但训练时间更长。结果表明,结合自然语言处理特征的传统分类器与深度学习模型均可有效区分新闻标题/链接的感知质量,存在预测性能与训练时间之间的权衡。

原文摘要 · Abstract (English)

The proliferation of online news enables potential widespread publication of perceived low-quality news headlines/links. As a result, we investigated whether it was possible to automatically distinguish perceived lower-quality news headlines/links from perceived higher-quality headlines/links. We evaluated twelve machine learning models on a binary, balanced dataset of 57,544,214 worldwide news website links/headings from 2018-2024 (28,772,107 per class) with 115 extracted linguistic features. Binary labels for each text were derived from scores based on expert consensus regarding the respective news domain quality. Traditional ensemble methods, particularly the bagging classifier, had strong performance (88.1% accuracy, 88.3% F1, 80/20 train/test split). Fine-tuned DistilBERT achieved the highest accuracy (90.3%, 80/20 train/test split) but required more training time. The results suggest that both NLP features with traditional classifiers and deep learning models can effectively differentiate perceived news headline/link quality, with some trade-off between predictive performance and train time.

新闻质量二分类NLPDistilBERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。