用自编码器量化词汇质量,更准确评估文本语言丰富性。
Autoencoder-Based Framework to Capture Vocabulary Quality in NLP
- 用神经网络容量代替传统指标衡量词汇丰富度
- 在欺诈文本和古籍数据集上验证方法有效
- 适合关注数据质量与模型适配的研究者
语言丰富性对自然语言处理至关重要,因数据特征常直接影响模型表现。然而,传统指标如类型-词频比(TTR)、词汇多样性(VOCD)和词汇文本多样性度量(MTLD)无法充分捕捉上下文关系、语义丰富性和结构复杂性。本文提出一种基于自编码器的框架,以神经网络容量作为词汇丰富度、多样性和复杂性的代理指标,实现对词汇规模、句法结构与上下文深度之间相互作用的动态评估。我们在两个不同数据集上验证该方法:涵盖多个领域欺骗性与欺诈性文本的DIFrauD数据集,以及涵盖多种语言、体裁和历史时期的Project Gutenberg数据集。实验结果表明该方法具有鲁棒性与适应性,为数据集筛选与NLP模型设计提供实用指导。通过提升传统词汇评估能力,本工作推动更具上下文感知与语言自适应能力的NLP系统发展。
原文摘要 · Abstract (English)
Linguistic richness is essential for advancing natural language processing (NLP), as dataset characteristics often directly influence model performance. However, traditional metrics such as Type-Token Ratio (TTR), Vocabulary Diversity (VOCD), and Measure of Lexical Text Diversity (MTLD) do not adequately capture contextual relationships, semantic richness, and structural complexity. In this paper, we introduce an autoencoder-based framework that uses neural network capacity as a proxy for vocabulary richness, diversity, and complexity, enabling a dynamic assessment of the interplay between vocabulary size, sentence structure, and contextual depth. We validate our approach on two distinct datasets: the DIFrauD dataset, which spans multiple domains of deceptive and fraudulent text, and the Project Gutenberg dataset, representing diverse languages, genres, and historical periods. Experimental results highlight the robustness and adaptability of our method, offering practical guidance for dataset curation and NLP model design. By enhancing traditional vocabulary evaluation, our work fosters the development of more context-aware, linguistically adaptive NLP systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。