用医学分类体系评估并减少视觉语言模型的误判风险。
Measuring and Aligning Abstraction in Vision-Language Models with Medical Taxonomies
- 引入层级指标和跨分支错误检测机制
- 严重抽象错误率降至2%以下
- 适合医疗AI安全评估与临床部署
视觉-语言模型在胸部X光分类任务中表现优异,但传统平铺指标无法区分临床轻重错误。本文通过医学分类体系,采用层级指标对多个先进VLM进行评估,提出“灾难性抽象错误”以捕捉跨分支误判。结果表明,尽管整体准确率高,现有模型与临床分类体系仍存在显著偏差。为此,我们提出风险约束阈值法和基于径向嵌入的分类体系感知微调,将严重抽象错误率降至2%以下,同时保持良好性能。研究强调了层级评估与表征对齐对医疗AI安全落地的重要性。
原文摘要 · Abstract (English)
Vision-Language Models show strong zero-shot performance for chest X-ray classification, but standard flat metrics fail to distinguish between clinically minor and severe errors. This work investigates how to quantify and mitigate abstraction errors by leveraging medical taxonomies. We benchmark several state-of-the-art VLMs using hierarchical metrics and introduce Catastrophic Abstraction Errors to capture cross-branch mistakes. Our results reveal substantial misalignment of VLMs with clinical taxonomies despite high flat performance. To address this, we propose risk-constrained thresholding and taxonomy-aware fine-tuning with radial embeddings, which reduce severe abstraction errors to below 2 per cent while maintaining competitive performance. These findings highlight the importance of hierarchical evaluation and representation-level alignment for safer and more clinically meaningful deployment of VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。