arXiv:2609.01909cs.AI2026-09

提出新框架,区分模型能力与数据测量的极限,帮医生判断该优化算法还是改进检测手段。

The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction

论文配图:The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction
图 1 · 摘自论文原文
  • 用学习者差距和测量天花板分离模型瓶颈与数据局限
  • 实测三组临床数据,发现多数模型仍可提升,但测量方式影响更大
  • 适合做医疗模型评估、临床研究设计或算法改进的开发者看

临床预测饱和可能源于模型无法提取信息,或数据变量本身存在上限。本文通过‘学习者差距’和‘测量通道天花板’分离二者,利用总变差分离刻画最优平衡准确率,实现架构无关性,给出替换污染下的精确部分识别结果,提出交叉拟合天花板估计器,并确定多模态决策改进的严格条件。新增两项有限样本诊断:标签置换乐观下限与欠拟合曲线。在三个真实队列中验证:UCI再入院(n=99,343)、BRFSS糖尿病(n=253,680)、NHANES HbA1c(n=10,219)。调优后的梯度提升几乎达到估计天花板,而低效模型仍有显著差距。NHANES显示问卷与实测边际天花板无差异,但联合使用有显著互补增益,反驳了‘客观模态必然主导’的简化观点。所有队列中,小的AUROC提升伴随大幅贝叶斯决策翻转率,不同架构估算相似天花板,但实际表现差异显著。对104项临床任务的PRISMA综述显示,通道级规律跨18种疾病重现:广泛但非普适的结构化临床区域,同通道增益递减,换测量通道则性能更高。该框架将饱和从经验观察变为可审计的决策依据:有余量时优化模型,无余量时改进测量。

原文摘要 · Abstract (English)

Clinical prediction can saturate for two different reasons: a fitted learner may fail to extract available information, or the recorded variables may impose a population frontier. We separate these quantities through the \emph{learner gap} and the \emph{measurement-channel ceiling}. Optimal balanced accuracy is characterized by total-variation separation, yielding architecture invariance, a sharp partial-identification result under replacement contamination, a cross-fitted ceiling estimator, and exact conditions for multimodal decision improvement. We add two finite-sample diagnostics, namely a label-permutation optimism floor and an underfit curve, and validate the audit on three real cohorts: UCI readmission ($n=99{,}343$), BRFSS diabetes ($n=253{,}680$), and NHANES HbA1c ($n=10{,}219$). Well-tuned gradient boosting nearly reaches the estimated frontier in UCI and BRFSS, whereas deliberately or practically deficient learners retain large gaps. NHANES yields a null difference between questionnaire and measured marginal frontiers but a significant joint complementarity gain, refining the simplistic claim that an objective modality must dominate. Across all cohorts, modest AUROC gains coexist with substantially larger Bayes decision-flip rates, and several architectures estimate similar frontiers while their achieved balanced accuracy differs sharply. A PRISMA-guided synthesis of 104 clinical tasks then shows that the same channel-level regularities recur across more than 18 disease categories: a broad but non-universal structured-clinical region, diminishing same-channel gains across model families, and higher performance when measurement channels change. The framework converts saturation from an empirical observation into an auditable decision: improve the learner when headroom remains; improve measurement when it does not.

临床预测模型评估医学人工智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。