arXiv:2512.00598cs.LGcs.AI2025-12被引 2

通过隐式分组提升脊柱手术并发症预测公平性

Developing Fairness-Aware Task Decomposition to Improve Equity in Post-Spinal Fusion Complication Prediction

  • 用聚类发现隐藏患者亚群,动态分配预测任务
  • 在四分类中达AUC 0.86,性别偏差降为0.055
  • 无需敏感属性也能实现可解释的公平预测

临床预测模型的公平性仍是重大挑战,尤其在青少年特发性脊柱侧弯矫形手术这类高风险场景中,患者预后差异显著。现有方法多依赖粗粒度人口学调整或事后修正,难以捕捉临床群体的潜在结构,甚至可能加剧偏见。本文提出FAIR-MTL框架,一种面向公平性的多任务学习方法,实现对术后并发症严重程度的精细、公平预测。该方法不直接使用敏感属性,而是通过数据驱动的方式推断潜在患者亚群:先提取紧凑的人口学嵌入,再以k-means聚类识别可能受传统模型影响不同的子群体,并据此决定共享多任务架构中的任务路由。训练时采用逆频加权缓解亚群不平衡,结合正则化防止对小群体过拟合。应用于四等级术后并发症预测,FAIR-MTL取得0.86的AUC和75%的准确率,优于单任务基线,同时显著降低偏差——性别方面,人口平等差降至0.055,等几率差降至0.094;年龄方面分别为0.056和0.148。模型可解释性通过SHAP与吉尼重要性分析保障,持续识别出血红蛋白、红细胞压积、体重等临床相关特征。结果表明,将无监督亚群发现融入多任务框架,能实现更公平、可解释且具临床价值的手术风险分层预测。

原文摘要 · Abstract (English)

Fairness in clinical prediction models remains a persistent challenge, particularly in high-stakes applications such as spinal fusion surgery for scoliosis, where patient outcomes exhibit substantial heterogeneity. Many existing fairness approaches rely on coarse demographic adjustments or post-hoc corrections, which fail to capture the latent structure of clinical populations and may unintentionally reinforce bias. We propose FAIR-MTL, a fairness-aware multitask learning framework designed to provide equitable and fine-grained prediction of postoperative complication severity. Instead of relying on explicit sensitive attributes during model training, FAIR-MTL employs a data-driven subgroup inference mechanism. We extract a compact demographic embedding, and apply k-means clustering to uncover latent patient subgroups that may be differentially affected by traditional models. These inferred subgroup labels determine task routing within a shared multitask architecture. During training, subgroup imbalance is mitigated through inverse-frequency weighting, and regularization prevents overfitting to smaller groups. Applied to postoperative complication prediction with four severity levels, FAIR-MTL achieves an AUC of 0.86 and an accuracy of 75%, outperforming single-task baselines while substantially reducing bias. For gender, the demographic parity difference decreases to 0.055 and equalized odds to 0.094; for age, these values reduce to 0.056 and 0.148, respectively. Model interpretability is ensured through SHAP and Gini importance analyses, which consistently highlight clinically meaningful predictors such as hemoglobin, hematocrit, and patient weight. Our findings show that incorporating unsupervised subgroup discovery into a multitask framework enables more equitable, interpretable, and clinically actionable predictions for surgical risk stratification.

医疗公平多任务学习聚类分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。