arXiv:2601.13710cs.LGcs.AI2026-01

用AI预测鼻窦手术效果,机器学习比生成式AI更准。

Who Benefits From Sinus Surgery? Comparing Generative AI and Supervised Machine Learning for Predicting Surgical Outcomes in Chronic Rhinosinusitis

  • 用临床数据训练机器学习模型预测术后改善情况。
  • 最佳模型准确率达85%,优于生成式AI的判别与校准能力。
  • 适合关注医疗决策支持和可解释性的临床研究者。

人工智能已重塑医学影像分析,但其在临床数据中的前瞻性决策支持应用仍有限。本文研究慢性鼻窦炎(CRS)患者术前对6个月后临床意义改善的预测,定义成功为SNOT-22评分下降超过8.9分(最小临床重要差异,MCID)。在一项前瞻性收集的全部患者均接受手术的队列中,我们评估仅使用术前临床数据的模型能否识别出可能预后不佳者(即应避免手术者)。对比监督学习(逻辑回归、树集成、自研MLP)与生成式AI(ChatGPT、Claude、Gemini、Perplexity),所有模型接收相同结构化输入,输出限制为二分类建议及置信度。最优监督学习模型(MLP)达到85%准确率,校准性与决策曲线净收益均更优。生成式AI在零样本设置下整体判别力与校准性表现较差。值得注意的是,生成式AI的推理理由与临床医生经验及MLP特征重要性一致,反复强调基线SNOT-22评分、CT/内镜严重程度、息肉表型以及心理/疼痛共病。研究提供可复现的表格到生成式AI评估协议与亚组分析。结果支持‘以机器学习为主、生成式AI为辅’的工作流:部署校准良好的机器学习模型进行手术适应症初筛,生成式AI作为解释工具提升透明度与共同决策能力。

原文摘要 · Abstract (English)

Artificial intelligence has reshaped medical imaging, yet the use of AI on clinical data for prospective decision support remains limited. We study pre-operative prediction of clinically meaningful improvement in chronic rhinosinusitis (CRS), defining success as a more than 8.9-point reduction in SNOT-22 at 6 months (MCID). In a prospectively collected cohort where all patients underwent surgery, we ask whether models using only pre-operative clinical data could have identified those who would have poor outcomes, i.e. those who should have avoided surgery. We benchmark supervised ML (logistic regression, tree ensembles, and an in-house MLP) against generative AI (ChatGPT, Claude, Gemini, Perplexity), giving each the same structured inputs and constraining outputs to binary recommendations with confidence. Our best ML model (MLP) achieves 85 % accuracy with superior calibration and decision-curve net benefit. GenAI models underperform on discrimination and calibration across zero-shot setting. Notably, GenAI justifications align with clinician heuristics and the MLP's feature importance, repeatedly highlighting baseline SNOT-22, CT/endoscopy severity, polyp phenotype, and physchology/pain comorbidities. We provide a reproducible tabular-to-GenAI evaluation protocol and subgroup analyses. Findings support an ML-first, GenAI- augmented workflow: deploy calibrated ML for primary triage of surgical candidacy, with GenAI as an explainer to enhance transparency and shared decision-making.

医疗AI预测模型生成式AI临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。