用大模型生成的文本嵌入可有效识别城市经济类型,比传统方法更清晰。
When Can We Work in Embedding Space? What Text Embeddings Preserve

- 基于主题混合生成模型,验证嵌入空间可替代原始文本
- 在363个都市区中,嵌入聚类能更精准分离就业动态
- 适合做城市经济分类或政策分析的研究者使用
文本嵌入能否作为经验分析的输入,依赖于一个假设:将文本替换为低维嵌入后损失甚微。本文在文档为潜在主题混合的生成模型下,使该假设具体化。研究了两类应用——在嵌入空间中聚类与控制高维文本。嵌入聚类对应具有相似主题混合的文档集合;控制嵌入等价于控制主题混合,有效性取决于该混合是否捕捉混淆因素。在363个美国都会区的应用中,基于大模型生成的经济描述嵌入进行聚类,恢复出可解释的经济原型,并比对模型残差聚类或基于行业与人口特征的精选协变量聚类,更清晰地区分本地就业动态。
原文摘要 · Abstract (English)
When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can trade text for its low-dimensional embedding, and lose little in doing so. I make that assumption precise under a generative model in which documents are mixtures of latent topics. I study two uses---clustering units in embedding space and controlling for high-dimensional text. A cluster of embeddings is a set of documents with similar topic mixtures; controlling for the embedding is equivalent to controlling for the topic mixture, so validity reduces to whether that mixture captures the confounding. In an application to 363 U.S. metropolitan areas, embedding-based clusters of LLM-generated economic descriptions recover interpretable economic archetypes and separate local employment dynamics more sharply than clustering on model residuals, or on a curated set of industry and demographic covariates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。