arXiv:2411.05192cs.CLcs.AI2024-11EMNLP被引 4

分析新闻中来源选择的写作规划,揭示记者如何根据主题选材。

Explaining Mixtures of Sources in News Articles

  • 基于记者实际写作习惯,构建8种来源选择框架
  • 发现立场和社交关联框架最能解释多数新闻来源安排
  • 仅凭标题即可较准确预测最适合的写作框架

人类写作者在动笔前会进行构思。为让大语言模型参与长文生成,需理解人类的构思过程。本文以新闻报道中的来源选择为例,研究记者如何规划使用不同来源。我们与专业记者合作,改进5种现有框架并提出3种新框架,用于描述新闻写作中来源整合的计划模式。受贝叶斯潜变量建模启发,设计指标以推断每篇报道背后最可能的框架。实验发现,立场和社交关联框架能最好解释大多数文章的来源安排;而在科学类等事实密集型话题中,文本蕴含框架更适用。此外,仅根据文章标题就可较准确预测最合适的框架。研究结果为理解人类写作规划提供了新范式,并发布包含400万篇文章标注的新闻来源数据集NewsSources。

原文摘要 · Abstract (English)

Human writers plan, then write. For large language models (LLMs) to play a role in longer-form article generation, we must understand the planning steps humans make before writing. We explore one kind of planning, source-selection in news, as a case-study for evaluating plans in long-form generation. We ask: why do specific stories call for specific kinds of sources? We imagine a generative process for story writing where a source-selection schema is first selected by a journalist, and then sources are chosen based on categories in that schema. Learning the article's plan means predicting the schema initially chosen by the journalist. Working with professional journalists, we adapt five existing schemata and introduce three new ones to describe journalistic plans for the inclusion of sources in documents. Then, inspired by Bayesian latent-variable modeling, we develop metrics to select the most likely plan, or schema, underlying a story, which we use to compare schemata. We find that two schemata: stance and social affiliation best explain source plans in most documents. However, other schemata like textual entailment explain source plans in factually rich topics like "Science". Finally, we find we can predict the most suitable schema given just the article's headline with reasonable accuracy. We see this as an important case-study for human planning, and provides a framework and approach for evaluating other kinds of plans. We release a corpora, NewsSources, with annotations for 4M articles.

新闻生成写作规划来源分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。