通过合并诊断码降低维度,提升住院成本模型的稳定性与可解释性
A Log-Linear Analytics Approach to Cost Model Regularization for Inpatient Stays through Diagnostic Code Merging
- 将ICD-10代码从7位缩至6位以下,减少特征维度
- 在保持可解释性的前提下,系数方差显著降低,模型更稳定
- 适合医疗成本建模、风险调整等需要长期一致性的场景
医疗研究中的成本模型需兼顾可解释性、准确性和参数一致性。然而,可解释模型常难以同时实现高精度与参数稳定。使用高度细粒度的ICD-10诊断码作为预测变量时,普通最小二乘法(OLS)虽具准确性,但回归系数随时间波动大,因多数编码在数据集中出现频率低。尽管正则化方法如Ridge可缓解此问题,却可能剔除关键预测因子。本文证明:降低ICD-10代码粒度是一种有效的OLS正则化策略,可在不丢失诊断类别信息的前提下,保持模型可解释性与一致性。通过将ICD-10代码从七位缩减至六位或更少,降低了回归问题的维度,数学上使海森矩阵的迹增大,从而减小系数估计的方差。研究结果解释了为何在实际风险调整和成本模型中,广义诊断分组(如DRGs、HCC码)优于细粒度的ICD-10码。
原文摘要 · Abstract (English)
Cost models in healthcare research must balance interpretability, accuracy, and parameter consistency. However, interpretable models often struggle to achieve both accuracy and consistency. Ordinary least squares (OLS) models for high-dimensional regression can be accurate but fail to produce stable regression coefficients over time when using highly granular ICD-10 diagnostic codes as predictors. This instability arises because many ICD-10 codes are infrequent in healthcare datasets. While regularization methods such as Ridge can address this issue, they risk discarding important predictors. Here, we demonstrate that reducing the granularity of ICD-10 codes is an effective regularization strategy within OLS while preserving the representation of all diagnostic code categories. By truncating ICD-10 codes from seven characters to six or fewer, we reduce the dimensionality of the regression problem while maintaining model interpretability and consistency. Mathematically, the merging of predictors in OLS leads to increased trace of the Hessian matrix, which reduces the variance of coefficient estimation. Our findings explain why broader diagnostic groupings like DRGs and HCC codes are favored over highly granular ICD-10 codes in real-world risk adjustment and cost models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。