让机器学习模型的医学阈值可审计、可解释、可调控。
Auditable Unit-Aware Thresholds in Symbolic Regression via Logistic-Gated Operators
- 用逻辑门控算子将阈值变为方程中的显式参数。
- 在多个临床数据集上恢复出符合临床实际的生理阈值。
- 适合医疗决策系统中需要可解释规则的场景。
AI 在医疗领域的规模化应用需兼顾准确性与可读性、可审计性、可管理性。临床与公共卫生决策常依赖数值阈值(如触发警报或治疗),但多数机器学习模型将这些阈值隐藏于不透明的评分或平滑响应曲线中。本文提出逻辑门控算子(LGO)用于符号回归,使阈值成为方程中的第一类参数,并映射回物理单位,便于与临床指南直接比对。在 MIMIC-IV ICU、eICU、NHANES 等公共重症监护与人群健康队列上,LGO 成功恢复了针对平均动脉压(MAP)、乳酸、格拉斯哥昏迷评分(GCS)、血氧饱和度(SpO2)、BMI、空腹血糖及腰围等指标的临床合理阈值,性能媲美 AutoScore 与可解释提升机(EBM)。阈值具有稀疏性和选择性:仅在数据支持分段切换时出现,平滑任务则自动剔除,生成简洁公式,便于临床医生审查、压力测试与修订。作为独立符号模型或黑箱系统的安全层,LGO 有助于将观察数据转化为可审计、单位明确的医疗规则,适用于其他阈值驱动领域。
原文摘要 · Abstract (English)
AI for health will only scale when models are not only accurate but also readable, auditable, and governable. Many clinical and public-health decisions hinge on numeric thresholds -- cut-points that trigger alarms, treatment, or follow-up -- yet most machine-learning systems bury those thresholds inside opaque scores or smooth response curves. We introduce logistic-gated operators (LGO) for symbolic regression, which promote thresholds to first-class, unit-aware parameters inside equations and map them back to physical units for direct comparison with guidelines. On public ICU and population-health cohorts (MIMIC-IV ICU, eICU, NHANES), LGO recovers clinically plausible gates on MAP, lactate, GCS, SpO2, BMI, fasting glucose, and waist circumference while remaining competitive with established scoring systems (AutoScore) and explainable boosting machines (EBM). The gates are sparse and selective: they appear when regime switching is supported by the data and are pruned on predominantly smooth tasks, yielding compact formulas that clinicians can inspect, stress-test, and revise. As a standalone symbolic model or a safety overlay on black-box systems, LGO helps translate observational data into auditable, unit-aware rules for medicine and other threshold-driven domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。