用图协方差和大模型分析分类特征交互,揭示数据背后的事件线索。
Explaining Categorical Feature Interactions Using Graph Covariance and LLMs
- 通过一热编码与图协方差捕捉分类特征随时间的依赖变化。
- 在CTDC数据集上识别出数千对显著互动特征及关键时间点。
- 结合大模型生成可解释的事件假说,适合数据洞察与社会问题研究者。
现代数据集常包含大量样本、丰富特征及时间戳,分析其背后事件需复杂统计方法与领域知识。本文聚焦全球反人口贩卖数据协作组织(CTDC)的合成数据集——涵盖2002至2022年超过20万条匿名记录,每条记录含众多分类特征。提出一种快速可扩展的方法,用于识别显著的分类特征交互,并利用大语言模型(LLMs)生成数据驱动的解释。方法首先对分类特征进行一热编码,随后在每个时间点计算图协方差,量化分类数据中依赖结构的时序变化;该度量在伯努利分布下具有一致性。基于此,识别出依赖性频繁变化或在特定时刻突然上升的特征对。这些特征对及其时间戳被输入LLM,生成潜在事件解释。通过大量模拟验证方法有效性,并应用于CTDC数据集,成功揭示有意义的特征对与背后的数据故事。
原文摘要 · Abstract (English)
Modern datasets often consist of numerous samples with abundant features and associated timestamps. Analyzing such datasets to uncover underlying events typically requires complex statistical methods and substantial domain expertise. A notable example, and the primary data focus of this paper, is the global synthetic dataset from the Counter Trafficking Data Collaborative (CTDC) -- a global hub of human trafficking data containing over 200,000 anonymized records spanning from 2002 to 2022, with numerous categorical features for each record. In this paper, we propose a fast and scalable method for analyzing and extracting significant categorical feature interactions, and querying large language models (LLMs) to generate data-driven insights that explain these interactions. Our approach begins with a binarization step for categorical features using one-hot encoding, followed by the computation of graph covariance at each time. This graph covariance quantifies temporal changes in dependence structures within categorical data and is established as a consistent dependence measure under the Bernoulli distribution. We use this measure to identify significant feature pairs, such as those with the most frequent trends over time or those exhibiting sudden spikes in dependence at specific moments. These extracted feature pairs, along with their timestamps, are subsequently passed to an LLM tasked with generating potential explanations of the underlying events driving these dependence changes. The effectiveness of our method is demonstrated through extensive simulations, and its application to the CTDC dataset reveals meaningful feature pairs and potential data stories underlying the observed feature interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。