从无结构数据中系统发现可解释的规律,避免人为偏见和误判。
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach
- 用概念嵌入将文本/音视频转为可测的高维特征
- 通过高维中心极限定理控制多重假设检验误差
- 自动生成自然语言描述,适合经济、社会科学研究者
社会科学家正越来越多地利用无结构数据(如文本、音频、视频)获取新实证发现,例如估算量化指标的描述性统计或因果效应。在许多场景下,研究者希望进行无监督探索——不预先指定关键变量,而是主动“发现”重要模式。本文提出一种通用且灵活的框架,以统计上严谨的方式实现这一目标:首先利用人工智能可解释性方法,将无结构数据点映射为高维、稀疏且可解释的“概念嵌入”;然后基于这些嵌入计算统计量,对每个概念逐个测试可解释的假设;接着使用经高维中心极限理论验证的算法进行选择性推断,生成被选中的“发现”集合;最后自动生成并评估人类可读的自然语言描述。该框架极大减少研究者自由度,有效抵御数据窥探和选择后推断问题,并支持快速、低成本的敏感性分析与复现。文中展示了其在实证经济学中对无结构数据的最新描述性与因果分析中的应用。
原文摘要 · Abstract (English)
Social scientists are increasingly turning to unstructured datasets to unlock new empirical insights, e.g., estimating descriptive statistics of or causal effects on quantitative measures derived from text, audio, or video data. In many settings, unsupervised analysis is of primary interest, in that the researcher does not want to (or cannot) manually pre-specify all important aspects of the unstructured data to measure; they are interested in "discovery." This paper proposes a general and flexible framework for pursuing such discovery from unstructured data in a statistically principled way. The framework leverages recent methods from the literature on AI interpretability to map unstructured data points to high-dimensional, sparse, and interpretable "concept embeddings"; computes statistics from these concept embeddings for testing interpretable, concept-by-concept hypotheses; performs selective inference on these hypotheses using algorithms validated by new results in high-dimensional central limit theory, producing a selected set ("discoveries"); and both generates and evaluates human-interpretable natural language descriptions of these discoveries. The proposed framework has few researcher degrees of freedom, is robust to data snooping and other post-selection inference concerns, and facilitates fast and inexpensive sensitivity analysis and replication. Applications to recent descriptive and causal analyses of unstructured data in empirical economics are explored.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。