用多源非线性测量重构隐变量回归,解决测量偏差导致的结论分歧问题。
Partial Identification with Multiple Nonlinear Measurements of a Latent Regressor

- 通过约束测量函数曲率差异,构建结构系数的闭式区间估计
- 四源以上可估计曲率边界,95%置信区间具均匀覆盖性
- 适用于职业暴露等多指标测量场景,适合需稳健结论的研究者
研究隐变量线性回归中,因观测变量仅通过多个噪声测量(每个为隐变量的光滑非线性函数)而产生的识别问题。该问题在人工智能职业暴露测量中尤为突出,不同评分方法导致下游估计相差十一倍。单一测量回归仅得特定来源系数,而非结构系数。通过要求共识测量函数为线性,并以斜率相对曲率异质性为界,推导出结构系数的闭式区间。该区间对未知来源载荷不变,半宽为曲率界值的二阶量级且紧致。至少四个测量时,可通过分裂工具辅助回归从联合分布估计曲率界,使用Imbens-Manski置信区间与Stoye临界值可实现对曲率类的统一覆盖,包括点识别边界。应用中将六种暴露度量匹配美国社区调查2015–2024年888万个人年观测数据。2022年后,语言模型测度与Webb专利文本测度的就业系数符号相反,事前因子分析规则将Webb测度识别为独立构念。保留五源后得出载荷不变的共识系数-0.239,部分识别半宽为点估计的1.23%,95%单侧上界下为1.88%。本应用视为测量校正,而非人工智能替代的因果估计。
原文摘要 · Abstract (English)
We study linear regression when the regressor is latent and observed only through multiple noisy measurements, each a smooth but possibly nonlinear function of the latent variable. The problem is acute in the measurement of occupational exposure to artificial intelligence, where competing scores yield downstream estimates that differ by a factor of eleven. A regression on any single measurement recovers a source-specific coefficient rather than the structural one. We fix the latent scale by requiring the consensus measurement function to be linear and bound the remaining curvature heterogeneity across sources relative to slope. Under this bound, the structural coefficient lies in a closed-form interval centered at a symmetric cross-source estimator. The interval is invariant to unknown source loadings, and its half-width is second order in the curvature bound and sharp to the same order. With at least four measurements, the bound is estimable from the joint distribution of the sources through a split-instrument auxiliary regression, and Imbens-Manski confidence intervals with the Stoye critical value attain uniform coverage over the curvature class, including at the point-identified boundary. The application matches six exposure measures to an American Community Survey panel of 8.88 million person-year observations for 2015 to 2024. The post-2022 employment coefficient changes sign between the language-model measures and the Webb patent-text measure, and an ex ante factor-analytic rule separates the Webb measure as a distinct construct. The five retained sources yield a loading-invariant consensus coefficient of -0.239, with a partial-identification half-width of 1.23 percent of the point estimate, or 1.88 percent at the one-sided 95 percent upper bound on the curvature. We read the application as measurement reconciliation rather than as a causal estimate of AI displacement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。