医学AI对错不只看准确率,还要看解释性、临床意义和责任分配。
What Does It Mean for a Medical AI System to Be Right?
- 从骨髓涂片分类出发,探讨AI正确性的多重维度
- 指出真实标签不稳定、模型过度自信、指标不适用等风险
- 适合关注AI医疗伦理与临床落地的研究者阅读
本文通过多发性骨髓瘤诊断中血浆细胞自动分类这一具体临床场景,探讨医学AI系统‘正确’的含义。结合科学哲学与研究伦理,论文认为医学AI的正确性并非单一可归约为基准性能的属性,而是涉及专家标注数据的可用性、模型输出的可解释性与可理解性、评估指标的临床意义,以及人机协作流程中的责任分配等多个维度。文章围绕四个相互关联的主题展开:真实标签的不稳定性、过度自信AI的不可见性、标准临床指标的局限性,以及在时间压力下自动化偏倚的风险。
原文摘要 · Abstract (English)
This paper examines what it means for a medical AI system to be right by grounding the question in a specific clinical context: the automatic classification of plasma cells in digitized bone marrow smears for the diagnosis of multiple myeloma. Drawing on philosophy of science and research ethics, the paper argues that correctness in medical AI is not a singular property reducible to benchmark performance, but a multi-dimensional concept involving the availability of expertly labeled medical datasets, the explainability and interpretability of model outputs, the clinical meaningfulness of evaluation metrics, and the distribution of accountability in human-AI workflows. As such, the paper develops this argument through four interrelated themes: the instability of ground truth labels, the opacity of overconfident AI, the inadequacy of standard clinical metrics, and the risk of automation bias in time-pressured clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。