arXiv:2501.00164cs.CLcs.CY2025-01

评测大模型识别新闻来源与理由的能力,为新闻真实性提供算法支持。

Measuring Large Language Models Capacity to Annotate Journalistic Sourcing

  • 构建五类来源标注框架,评估模型识别新闻来源信号的能力
  • 模型准确率较低,尤其在识别来源理由上表现较差
  • 为新闻伦理自动化评估提供首个系统性基准方案,适合媒体技术研究者

自2022年底ChatGPT发布以来,大语言模型的能力及其评估持续引发学术界与产业界的讨论。尽管法律、医学、数学等领域已建立多种场景与基准(Bommasani et al., 2023),但新闻领域尤其是新闻来源与伦理问题仍缺乏充分的评估场景。新闻是民主社会事实判断的关键功能(Vincent, 2023),而来源是原创新闻的核心支柱。评估大模型对新闻故事中来源信号及其记者论证依据的标注能力,对于构建更透明、更符合伦理的新闻自动化系统具有重要意义。本文提出一个基于新闻学研究(Gans, 2004)的五分类标注框架,设计使用案例、数据集与评估指标,作为系统性基准化的第一步。实验结果表明,当前大模型在识别故事中的所有来源陈述方面仍有较大提升空间,对来源类型匹配也存在困难,而识别来源正当性理由更是挑战极大。

原文摘要 · Abstract (English)

Since the launch of ChatGPT in late 2022, the capacities of Large Language Models and their evaluation have been in constant discussion and evaluation both in academic research and in the industry. Scenarios and benchmarks have been developed in several areas such as law, medicine and math (Bommasani et al., 2023) and there is continuous evaluation of model variants. One area that has not received sufficient scenario development attention is journalism, and in particular journalistic sourcing and ethics. Journalism is a crucial truth-determination function in democracy (Vincent, 2023), and sourcing is a crucial pillar to all original journalistic output. Evaluating the capacities of LLMs to annotate stories for the different signals of sourcing and how reporters justify them is a crucial scenario that warrants a benchmark approach. It offers potential to build automated systems to contrast more transparent and ethically rigorous forms of journalism with everyday fare. In this paper we lay out a scenario to evaluate LLM performance on identifying and annotating sourcing in news stories on a five-category schema inspired from journalism studies (Gans, 2004). We offer the use case, our dataset and metrics and as the first step towards systematic benchmarking. Our accuracy findings indicate LLM-based approaches have more catching to do in identifying all the sourced statements in a story, and equally, in matching the type of sources. An even harder task is spotting source justifications.

新闻生成大模型评估来源识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。