arXiv:2507.01764cs.CL2025-07被引 1

emoji和同形字影响文本分词,破坏语料库准确性

Data interference: emojis, homoglyphs, and issues of data fidelity in corpora and their results

  • 分析表情符号与同形字对分词的影响机制
  • 发现未处理这些元素会扭曲语言数据表示
  • 适合语料库语言学与数字人文研究者参考

分词是语料库语言学的关键步骤,为定量方法(如共现词分析)提供基础,并确保定性分析的可靠性。本文探讨分词差异如何影响语言数据的表征与分析结果的有效性,重点关注表情符号与同形字带来的挑战。研究表明,需对这些元素进行预处理以维持语料库对源数据的真实性。论文提出保障数字文本在语料库中准确表征的方法,从而支持可靠的语言分析并确保研究可重复性。研究强调需同时理解语言与技术层面,以提升语料分析精度,对基于语料的定量与定性研究均有重要意义。

原文摘要 · Abstract (English)

Tokenisation - "the process of splitting text into atomic parts" (Brezina & Timperley, 2017: 1) - is a crucial step for corpus linguistics, as it provides the basis for any applicable quantitative method (e.g. collocations) while ensuring the reliability of qualitative approaches. This paper examines how discrepancies in tokenisation affect the representation of language data and the validity of analytical findings: investigating the challenges posed by emojis and homoglyphs, the study highlights the necessity of preprocessing these elements to maintain corpus fidelity to the source data. The research presents methods for ensuring that digital texts are accurately represented in corpora, thereby supporting reliable linguistic analysis and guaranteeing the repeatability of linguistic interpretations. The findings emphasise the necessity of a detailed understanding of both linguistic and technical aspects involved in digital textual data to enhance the accuracy of corpus analysis, and have significant implications for both quantitative and qualitative approaches in corpus-based research.

语料库语言学分词数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。