arXiv:2503.10470cs.CL2025-03

用ASCII码和主成分分析法,无须语法标注就能评估句子结构平衡。

Statistical Analysis of Sentence Structures through ASCII, Lexical Alignment and PCA

  • 基于ASCII码与压缩后词性对齐的主成分分析,无需语法标注。
  • 11个语料中4个通过正态性检验,大模型输出接近正态分布。
  • 适合资源受限场景下的文本质量与风格分析,可补充传统语法工具。

尽管词性标注等句法工具有助于理解句子结构及其在不同语料中的分布,但其复杂性给自然语言处理带来了挑战。本研究聚焦于不依赖此类工具,通过美国信息交换标准代码(ASCII)表示11个来自不同来源的文本语料,并对其压缩后的词汇类别进行对齐,利用主成分分析(PCA)处理,再通过直方图与正态性检验(如Shapiro-Wilk和Anderson-Darling检验)分析结果。该方法通过关注ASCII码简化了文本处理流程,虽不替代句法工具,但作为资源高效的辅助手段,可用于评估文本结构平衡。由Grok生成的文本表现出接近正态分布,表明大模型输出具有良好的句法平衡性;其余10个语料中有4个通过正态性检验。未来研究可探索其在文本质量评估与风格分析中的应用,结合句法信息拓展更广泛任务。

原文摘要 · Abstract (English)

While utilizing syntactic tools such as parts-of-speech (POS) tagging has helped us understand sentence structures and their distribution across diverse corpora, it is quite complex and poses a challenge in natural language processing (NLP). This study focuses on understanding sentence structure balance - usages of nouns, verbs, determiners, etc - harmoniously without relying on such tools. It proposes a novel statistical method that uses American Standard Code for Information Interchange (ASCII) codes to represent text of 11 text corpora from various sources and their lexical category alignment after using their compressed versions through PCA, and analyzes the results through histograms and normality tests such as Shapiro-Wilk and Anderson-Darling Tests. By focusing on ASCII codes, this approach simplifies text processing, although not replacing any syntactic tools but complementing them by offering it as a resource-efficient tool for assessing text balance. The story generated by Grok shows near normality indicating balanced sentence structures in LLM outputs, whereas 4 out of the remaining 10 pass the normality tests. Further research could explore potential applications in text quality evaluation and style analysis with syntactic integration for more broader tasks.

文本分析统计方法大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。