构建多源视频可复现的肺部超声评估基准,检验AI模型跨域泛化与任务有效性。
External Benchmarking of Lung Ultrasound Models for Pneumothorax-Related Signs: A Manifest-Based Multi-Source Study
- 基于190个公开视频构建带标签的视频片段清单,支持外部复现评估
- 单一机构模型在外部数据上AUC降至0.705,跨域性能显著下降
- 发现肺搏动与肺滑动缺失差异明显,揭示二分类模型忽视临床中间状态
目前针对气胸相关肺超声(LUS)AI模型的可复现外部评估稀缺,且二分类肺滑动判断可能掩盖重要临床征象。为此,我们构建了一个基于清单的外部评估基准,用于检验跨领域泛化能力与任务有效性。从190个公开可获取的LUS视频中收集280段视频片段,并发布包含视频链接、时间戳、裁剪坐标、标签和探头形状的重建清单。标签包括正常肺滑动、无肺滑动、肺点、肺搏动。将此前发表的单中心二分类模型在此基准上进行评估;并通过预测无肺滑动概率P(absent)分析肺点与肺搏动的表现。结果显示,该模型在域内ROC-AUC为0.9625,但在异构外部基准上降至0.7050;即使仅限线性片段评估,仍为0.7212。挑战态分析显示,均值P(absent)排序为:无肺滑动(0.504)> 肺点(0.313)> 正常(0.186)> 肺搏动(0.143)。肺搏动与无肺滑动差异显著(p=0.000470),但与正常无差异(p=0.813),表明该二分类模型将肺搏动视为类似正常;而肺点与两者均有显著差异(分别p=0.000468, p=0.000026),支持其作为介于二者之间的模糊状态而非清晰二元类别。结论:基于清单的多源基准可实现无需重传原始视频的可复现外部评估。二分类肺滑动判断是气胸推理的不完整代理,因其掩盖了盲区与模糊状态如肺搏动与肺点。
原文摘要 · Abstract (English)
Background and Aims: Reproducible external benchmarks for pneumothorax-related lung ultrasound (LUS) AI are scarce, and binary lung-sliding classification may obscure clinically important signs. We therefore developed a manifest-based external benchmark and used it to test both cross-domain generalization and task validity. Methods: We curated 280 clips from 190 publicly accessible LUS source videos and released a reconstruction manifest containing URLs, timestamps, crop coordinates, labels, and probe shape. Labels were normal lung sliding, absent lung sliding, lung point, and lung pulse. A previously published single-site binary classifier was evaluated on this benchmark; challenge-state analysis examined lung point and lung pulse using the predicted probability of absent sliding, P(absent). Results: The single-site comparator achieved ROC-AUC 0.9625 in-domain but 0.7050 on the heterogeneous external benchmark; restricting external evaluation to linear clips still yielded ROC-AUC 0.7212. In challenge-state analysis, mean P(absent) ranked absent (0.504) > lung point (0.313) > normal (0.186) > lung pulse (0.143). Lung pulse differed from absent clips (p=0.000470) but not from normal clips (p=0.813), indicating that the binary model treated pulse as normal-like despite absent sliding. Lung point differed from both absent (p=0.000468) and normal (p=0.000026), supporting its interpretation as an intermediate ambiguity state rather than a clean binary class. Conclusion: A manifest-based, multi-source benchmark can support reproducible external evaluation without redistributing source videos. Binary lung-sliding classification is an incomplete proxy for pneumothorax reasoning because it obscures blind-spot and ambiguity states such as lung pulse and lung point.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。