监控西班牙媒体英语借词使用,构建动态数据库
Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press

- 用神经序列标注模型自动识别西班牙语新闻中的英语借词
- 发现平均每千词出现2个未融合借词,六年间持续增长
- 适合语言政策研究者、媒体分析者及跨语言传播研究者
本文介绍Observatorio Lázaro,一个监测西班牙数字媒体中未融合词汇借用(主要为英语借词)的语言资源。自2020年4月起,系统每日自动处理多家新闻机构的稿件,通过神经序列标注模型检测借词,并通过公共网页界面和API提供结果。截至撰写时,数据库已收录超过两百万次借词,覆盖188万篇文章与9.93亿词元文本(2020–2026)。论文详述了端到端流程(采集、检测、后处理、存储与访问)、数据模型及开放条款;评估显示检测器在保留测试集上对借词类别的跨度级F1为0.86,训练语料标注者间一致性Cohen's kappa达0.91,人工抽检1000个片段精确率达93%。统计分析表明,西班牙语借词频率约每千词2次,保持稳定;其词汇呈开放增长趋势,58.7%的词形仅出现一次(校正后为53.6%),在时尚、科技与生活方式板块密度最高,政治与制度类新闻最低。该资源旨在补充静态词典与一次性标注语料,提供持续更新的媒体借词记录。
原文摘要 · Abstract (English)
This paper describes Observatorio Lázaro, a language resource that monitors unassimilated lexical borrowings (predominantly English lexical borrowings or anglicisms) in the Spanish digital press. Since April 2020 the system has automatically processed the daily output of a collection of news outlets, detected borrowings with a neural sequence-labeling model, and made the results available through a public web interface and API. The result is a continuously updated diachronic database which, at the time of writing, records more than two million borrowings across 1.88 million articles and 993 million running tokens of text (2020-2026). The paper documents the resource: we describe the end-to-end pipeline (acquisition, detection, post-processing, storage and access), the data model and the terms of availability; we evaluate the resource through the detector's held-out performance (span-level F1=0.86 for the borrowing class), inter-annotator agreement on the training corpus (Cohen's kappa=0.91) and a manual precision audit of 1,000 spans from the deployed data; and we situate it with respect to Spanish borrowing lexicography, annotated borrowing corpora and neology-monitoring observatories. The data shows that unassimilated anglicisms are used in the Spanish press at a frequency of approximately two anglicisms per thousand tokens, and that this rate remains stable. Our statistical analysis over six years reveals that the anglicism vocabulary in Spanish behaves as an open and growing class, with 58.7% of its types attested only once (53.6% after correcting for detection precision), and that its density is highest in the fashion, technology and lifestyle sections and lowest in political and institutional news. The resource is intended to complement static borrowing dictionaries and one-off annotated corpora by providing a continuously updated record of borrowing in the Spanish press.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。