研究短周期定价数据中因设计稀疏导致的推断误差,提出新方法提升统计可靠性。
Across-Design Uncertainty in Short Pricing Panels: Inference and Identification

- 区分实际价格路径与不同设计路径间的不确定性,揭示主要误差来源
- 跨设计方差占总误差97.6%,标准方法无法捕捉此偏差
- 通过区域随机化生成独立识别变异可显著改善推断覆盖度
短周期观测定价面板通常样本量大但价格变动少。我们通过模拟数据生成过程,将估计误差分解为给定价格轨迹下的不确定性与不同轨迹间的变异性。基准模拟显示,对梯度提升模型而言,跨设计成分占估计误差方差的97.6%,导致覆盖率不足,这是由设计相关的中心化误差引起的,而传统的组内重抽样和聚类稳健方法未能捕捉该问题。三个核心发现:第一,跨设计离散度满足经验关系 sigma_b ≈ 0.182 V^(-0.271),其中 V = n_moves × magnitude²,指数-0.271视为模拟规律;第二,共享相同价格路径的区域虽能改进干扰项估计,但不产生独立价格轨迹;唯有在具有独立设计误差的单位间平均,才能以√k速率降低跨设计标准差;第三,基于独立定价单位的Paule-Mandel方差分量估计,在同质性假设下使实证覆盖率从0.469提升至0.931。总体而言,提升被动面板推断能力需主动生成独立识别变异,例如通过受控区域随机化。一项针对扫描数据(Dominick's Finer Foods, Soft Drinks)的应用证实:名义价格区与产品在独立设计抽样中仅表现为极小部分,导致单位间分散区间远宽于传统组内自助法。
原文摘要 · Abstract (English)
Short observational pricing panels often contain many observations but few distinct price movements. We evaluate the inferential consequences of this sparsity in a synthetic data-generating process by separating estimation error into uncertainty conditional on a realized price trajectory and variation across alternative trajectories. In baseline simulations, this across-design component accounts for 97.6% of estimation error variance for a gradient-boosted specification, causing coverage shortfalls driven by design-specific centering error that standard within-panel resampling and cluster-robust procedures fail to capture. Three main results organize the analysis. First, across-design dispersion follows the empirical relation sigma_b approx 0.182 V^(-0.271), where V = n_moves * magnitude^2, with -0.271 treated as a simulation regularity. Second, adding regions sharing a common price path improves nuisance estimation but creates no independent price trajectories; only averaging across units with independent design errors reduces across-design standard deviation at the sqrt(k) rate. Third, a Paule-Mandel variance component estimated across independently priced units increases empirical coverage under homogeneity from 0.469 to 0.931. Broadly, improving inference in passive panels requires generating independent identifying variation, such as through controlled regional randomization. Finally, an application to scanner data (Dominick's Finer Foods, Soft Drinks) confirms these findings: nominal price zones and products behave as a small fraction of their count in independent design draws, yielding between-unit dispersion intervals far wider than conventional within-panel bootstraps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。