arXiv:2411.00964cs.CL2024-11

用预训练词向量自动生成可解释的文本评分工具

Generic Embedding-Based Lexicons for Transparent and Reproducible Text Scoring

  • 基于FastText和GloVe词向量构建词典,仅需少量人工干预
  • 生成的词典在透明性和性能上优于传统手工词典
  • 适合需要可复现、可解释文本分析的研究者使用

过去十年间,文本分析工具日益复杂,研究者面临两难:使用高性能但不透明且计算开销大的先进模型,或依赖透明易用但性能有限的手工词典。本文提出一种折中方案:利用通用预训练词向量(FastText和GloVe 6B)自动生成词典,仅需少量人工输入。通过构建一系列概念性词典,证明嵌入式词典能够满足对透明性与高性能兼具的文本测量需求。

原文摘要 · Abstract (English)

With text analysis tools becoming increasingly sophisticated over the last decade, researchers now face a decision of whether to use state-of-the-art models that provide high performance but that can be highly opaque in their operations and computationally intensive to run. The alternative, frequently, is to rely on older, manually crafted textual scoring tools that are transparently and easily applied, but can suffer from limited performance. I present an alternative that combines the strengths of both: lexicons created with minimal researcher inputs from generic (pretrained) word embeddings. Presenting a number of conceptual lexicons produced from FastText and GloVe (6B) vector representations of words, I argue that embedding-based lexicons respond to a need for transparent yet high-performance text measuring tools.

词典生成可解释性文本分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。