构建多语言语料库,支持社会科学新概念研究。
Multilingual corpora for the study of new concepts in the social sciences and humanities:
- 融合企业官网与年报数据,自动提取并清洗文本。
- 每条术语上下文提取5句,标注主题类别用于分类训练。
- 可复现扩展,适合研究新兴概念与NLP应用。
本文提出一种混合方法,构建多语言语料库以支持人文学科和社会科学中新兴概念的研究,以“非技术性创新”为例。语料库基于两类互补来源:(1)从公司网站自动提取并清洗的法语和英语文本;(2)按年份、格式和重复性标准自动筛选的年度报告。处理流程包括自动语言识别、无关内容过滤、相关片段提取及结构化元数据补充。从初始语料库中构建英文衍生数据集,用于机器学习。针对专家词表中的每个术语,提取其所在句前后各两句话(共五句)作为上下文块,并标注对应主题类别,形成适用于监督分类任务的数据。该方法生成可复现、可扩展的资源,既可用于分析新兴概念的词汇多样性,也可为自然语言处理提供训练数据。
原文摘要 · Abstract (English)
This article presents a hybrid methodology for building a multilingual corpus designed to support the study of emerging concepts in the humanities and social sciences (HSS), illustrated here through the case of ``non-technological innovation''. The corpus relies on two complementary sources: (1) textual content automatically extracted from company websites, cleaned for French and English, and (2) annual reports collected and automatically filtered according to documentary criteria (year, format, duplication). The processing pipeline includes automatic language detection, filtering of non-relevant content, extraction of relevant segments, and enrichment with structural metadata. From this initial corpus, a derived dataset in English is created for machine learning purposes. For each occurrence of a term from the expert lexicon, a contextual block of five sentences is extracted (two preceding and two following the sentence containing the term). Each occurrence is annotated with the thematic category associated with the term, enabling the construction of data suitable for supervised classification tasks. This approach results in a reproducible and extensible resource, suitable both for analyzing lexical variability around emerging concepts and for generating datasets dedicated to natural language processing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。