arXiv:2604.05482cs.CVcs.AI2026-04

用视觉语言模型指导生成分割,结合统计理论检测犬类气胸,结果可解释且精准。

Unifying VLM-Guided Flow Matching and Spectral Anomaly Detection for Interpretable Veterinary Diagnosis

  • 用VLM引导流匹配迭代优化肺部病灶分割边界。
  • 通过随机矩阵理论识别异常特征值,准确率超90%。
  • 适合需要高可信度医疗诊断的临床与研究场景。

犬类气胸自动诊断面临数据稀缺与模型可信度挑战。为此,我们首次发布一个公开的像素级标注数据集以促进研究。提出一种新型诊断范式,将任务重构为信号定位与谱检测的协同过程。定位阶段,采用视觉语言模型(VLM)引导迭代流匹配,逐步精炼分割掩码,实现高精度边界。检测阶段,利用分割掩码提取可疑病灶特征,应用随机矩阵理论(RMT)分析这些特征。该方法将健康组织建模为可预测的随机噪声,通过检测显著偏离的特征值来识别非随机病理信号。流匹配提供的高保真定位有效净化信号,极大提升RMT检测器敏感性。生成分割与第一性原理统计分析的协同,构建了高精度且可解释的诊断系统(源代码见:https://github.com/Pu-Wang-alt/Canine-pneumothorax)。

原文摘要 · Abstract (English)

Automatic diagnosis of canine pneumothorax is challenged by data scarcity and the need for trustworthy models. To address this, we first introduce a public, pixel-level annotated dataset to facilitate research. We then propose a novel diagnostic paradigm that reframes the task as a synergistic process of signal localization and spectral detection. For localization, our method employs a Vision-Language Model (VLM) to guide an iterative Flow Matching process, which progressively refines segmentation masks to achieve superior boundary accuracy. For detection, the segmented mask is used to isolate features from the suspected lesion. We then apply Random Matrix Theory (RMT), a departure from traditional classifiers, to analyze these features. This approach models healthy tissue as predictable random noise and identifies pneumothorax by detecting statistically significant outlier eigenvalues that represent a non-random pathological signal. The high-fidelity localization from Flow Matching is crucial for purifying the signal, thus maximizing the sensitivity of our RMT detector. This synergy of generative segmentation and first-principles statistical analysis yields a highly accurate and interpretable diagnostic system (source code is available at: https://github.com/Pu-Wang-alt/Canine-pneumothorax).

医学影像可解释性异常检测生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。