arXiv:2502.03252cs.CL2025-02被引 1

构建概念口语与书面语的量化尺度,实现文本自动分类。

A scale of conceptual orality and literacy: Automatic text categorization in the tradition of "Nähe und Distanz"

  • 基于主成分分析建立概念口语/书面语量化尺度
  • 区分口语与书面特征可更精细地排序文本
  • 适用于语料库编纂与大规模分析的辅助工具

Koch和Oesterreicher提出的“Nähe und Distanz”理论(Nähe=即时性,概念口语;Distanz=距离感,概念书面)在德语语言学中广泛应用。然而,该理论缺乏统计基础,难以融入实证语料库研究。本文基于主成分分析(PCA)构建了概念口语与书面语的量化尺度,并结合自动分析方法,以新高地德语两个语料库为例验证。研究发现,必须区分概念口语与书面特征,才能实现文本的精细化排序。该尺度可用于语料库建设及大规模分析中的引导与控制,其理论驱动的特性使其在功能上优于Biber的维度1。

原文摘要 · Abstract (English)

Koch and Oesterreicher's model of "Nähe und Distanz" (Nähe = immediacy, conceptual orality; Distanz = distance, conceptual literacy) is constantly used in German linguistics. However, there is no statistical foundation for use in corpus linguistic analyzes, while it is increasingly moving into empirical corpus linguistics. Theoretically, it is stipulated, among other things, that written texts can be rated on a scale of conceptual orality and literacy by linguistic features. This article establishes such a scale based on PCA and combines it with automatic analysis. Two corpora of New High German serve as examples. When evaluating established features, a central finding is that features of conceptual orality and literacy must be distinguished in order to rank texts in a differentiated manner. The scale is also discussed with a view to its use in corpus compilation and as a guide for analyzes in larger corpora. With a theory-driven starting point and as a "tailored" dimension, the approach compared to Biber's Dimension 1 is particularly suitable for these supporting, controlling tasks.

语料库语言学文本分类量化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。