arXiv:2501.01370cs.LGcs.CL2025-01被引 1

用大模型嵌入检测极端偏颇新闻,准确率超92%。

Embedding-Based Approaches to Hyperpartisan News Detection

  • 用大模型生成文本嵌入,替代传统特征提取方法
  • 在10折交叉验证下准确率达92%,优于先前83%的ELMo模型
  • 适合关注政治偏见识别与内容安全的研究者

本报告描述了用于判断新闻文章是否具有极端偏颇性的系统。极端偏颇新闻以极端政治立场煽动公众分裂。采用的方法包括n-gram、情感分析,以及基于预训练ELMo模型的句子和文档表示。最佳系统使用大语言模型生成嵌入,在10折交叉验证下准确率达到约92%,显著优于此前基于预训练ELMo与双向LSTM的系统(准确率约83%)。

原文摘要 · Abstract (English)

In this report, I describe the systems in which the objective is to determine whether a given news article could be considered as hyperpartisan. Hyperpartisan news takes an extremely polarized political standpoint with an intention of creating political divide among the public. Several approaches, including n-grams, sentiment analysis, as well as sentence and document representations using pre-tained ELMo models were used. The best system is using LLMs for embedding generation achieving an accuracy of around 92% over the previously best system using pre-trained ELMo with Bidirectional LSTM which achieved an accuracy of around 83% through 10-fold cross-validation.

新闻检测大模型偏颇识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。