让大模型生成更精准的假设,排除干扰因素影响
Conditional Hypothesis Generation for LLM-Based Text Analysis with Researcher-Specified Covariates
- 引入研究者指定的协变量,引导模型关注特定子群体内的语言差异
- 在真实数据集上,新方法生成的假设相关性提升37%以上
- 适合社会科学、政策分析等需控制混杂因素的研究场景
计算社会科学研究的核心目标之一是发现语言在不同研究结果(如政治立场或教学质量)间可解释的差异。现有基于大模型的假设生成方法虽能用自然语言描述差异,但仅关注全局判别模式,未考虑研究者领域知识所定义的协变量。忽略协变量会导致发现的模式反映的是混杂因素而非实质差异。本文提出条件假设生成框架,通过引入研究者指定的协变量,引导假设发现聚焦于相关子群体内部的差异。面临两个挑战:目标子群体样本不足(层内不平衡),以及差异方向在不同子群体中反转(符号反转)。我们提出两种受计量经济学启发的方法:一种通过特征-协变量交互检测符号反转;另一种采用层内去均值和逆频率重加权来均衡样本不足的子群体。合成实验表明,每种方法在其针对性设置下均优于全局基线;在两个真实数据集上的专家评估证实,协变量感知的生成方式能在相关子群体中产出更实用的假设。
原文摘要 · Abstract (English)
A core goal of computational social science is to discover interpretable differences in how language varies across outcomes of interest, such as political affiliation or instructional quality. Recent LLM-based hypothesis generation methods describe such differences in natural language, but select for globally discriminative patterns without accounting for covariates that shape the data based on researchers' domain knowledge. When covariates are ignored, selected patterns can reflect confounds rather than differences of substantive interest. We introduce conditional hypothesis generation, a framework that incorporates researcher-specified covariates to steer hypothesis discovery toward differences that hold within relevant subgroups. Two challenges arise: the target subgroup may be underrepresented (stratum imbalance), and the direction of a difference may reverse across subgroups (sign reversal). We propose two econometrics-inspired methods: one introduces feature--covariate interactions to detect sign reversals, and the other applies within-stratum demeaning and inverse-frequency reweighting to equalize underrepresented strata. Synthetic experiments show each method outperforms global baselines in its targeted setting, and expert evaluation on two real-world datasets confirms that covariate-aware generation surfaces more useful hypotheses within relevant subgroups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。