发现新闻机构间跨语言内容重用现象,定位重用位置并支持记者减负。
Rewrite the News: Tracing Editorial Reuse Across News Agencies
- 基于发布时间推断最早来源,无须完整翻译即可检测跨语言句子重用。
- 52%的斯洛文尼亚新闻稿含重用内容,多出现在文章中后段且以改写为主。
- 方法适用于多语种新闻分析,助力记者快速识别重复报道来源。
本文研究多语言新闻中的句级文本重用现象,分析重用内容在文章中的分布位置。提出一种弱监督方法,在无需完整翻译的情况下检测跨语言句级重用,旨在为记者提供自动化预筛选,缓解信息过载问题(Holyst et al., 2024)。研究对比了斯洛文尼亚新闻社(STA)的英文报道与15家外文机构(FA)在七种语言中的报告,利用发布时戳锁定每句重用内容的最早可能来源。分析涵盖2023年10月7日至11月2日、2025年2月1日至28日两个时间段,共1,037篇STA文章和237,551篇FA文章,筛选出1,087对对齐句子。结果显示,52%的STA文章存在重用,而FA文章仅1.6%有重用;重用以非字面形式为主,包含改写及来自多个来源的组合式重用。重用内容多集中于英文文章的中后段,导语部分更常为原创,说明单纯词汇匹配会遗漏大量编辑性重用。相比以往专注单语重叠的研究,本工作首次实现:(i) 无需完整翻译的跨语言重用检测,(ii) 利用发布时间确定潜在来源,(iii) 分析重用内容在文章中的具体位置。数据与代码已公开:https://github.com/kunturs/lrec2026-rewrite-news。
原文摘要 · Abstract (English)
This paper investigates sentence-level text reuse in multilingual journalism, analyzing where reused content occurs within articles. We present a weakly supervised method for detecting sentence-level cross-lingual reuse without requiring full translations, designed to support automated pre-selection to reduce information overload for journalists (Holyst et al., 2024). The study compares English-language articles from the Slovenian Press Agency (STA) with reports from 15 foreign agencies (FA) in seven languages, using publication timestamps to retain the earliest likely foreign source for each reused sentence. We analyze 1,037 STA and 237,551 FA articles from two time windows (October 7-November 2, 2023; February 1-28, 2025) and identify 1,087 aligned sentence pairs after filtering to the earliest sources. Reuse occurs in 52% of STA articles and 1.6% of FA articles and is predominantly non-literal, involving paraphrase and compositional reuse from multiple sources. Reused content tends to appear in the middle and end of English articles, while leads are more often original, indicating that simple lexical matching overlooks substantial editorial reuse. Compared with prior work focused on monolingual overlap, we (i) detect reuse across languages without requiring full translation, (ii) use publication timing to identify likely sources, and (iii) analyze where reused material is situated within articles. Dataset and code: https://github.com/kunturs/lrec2026-rewrite-news.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。