通过分解面部图像特征提升睡眠呼吸暂停筛查准确率
Structured Visual Evidence Decomposition for Evidence-Grounded Multimodal Screening of Obstructive Sleep Apnea-Hypopnea Syndrome

- 将面部图像拆解为7个解剖区域提问,生成结构化证据卡
- 在642人数据集上实现94.86%敏感度与5.14%假阴性率
- 适合医疗筛查场景,可作为辅助诊断的可审计工具
针对阻塞性睡眠呼吸暂停低通气综合征(OSAHS)的有效术前筛查,需结合临床风险因素与可见的头颈部解剖特征。直接调用通用多模态大模型进行医学二分类决策,易导致输出不稳定、校准不足。本文提出EviOSAHS,一种基于证据的多模态推理框架,将图像仅解剖证据获取与最终临床判断分离。每张正面面部图像被分解为七个固定解剖查询:颈部、下巴、口部、面颈脂肪、下颌、中面部和鼻部。视觉响应转化为结构化证据卡,记录目标解剖区域、可见性、风险方向、证据强度、置信度及摘要。这些卡片仅在最后阶段与清洗后的临床资料融合,由大语言模型执行平衡的二分类筛查判断。在642名受试者队列上评估,正常者归为筛查阴性,轻、中、重度OSAHS者归为筛查阳性。EviOSAHS达到88.47%准确率、94.86%敏感度、93.74% F1分数,假阴性率为5.14%,优于临床提示、直接多模态提示及朴素两阶段流程。消融实验表明,七问视觉分解与平衡最终判断对高敏感度至关重要。4,494次视觉输出的逐题审计显示,结构化解析率达100%,高可见性率达93.88%。EviOSAHS提供了一种可审计、高敏感度的二元术前筛查工作流,但应作为分诊助手而非诊断系统使用。临床部署前仍需前瞻性验证、外部测试及校准操作点控制。
原文摘要 · Abstract (English)
Effective pre-polysomnography screening for obstructive sleep apnea-hypopnea syndrome (OSAHS) requires combining clinical risk factors with visible craniofacial and neck cues. Directly prompting general-purpose multimodal foundation models for medical yes/no decisions can yield unstable, poorly calibrated outputs. We propose EviOSAHS, an evidence-grounded multimodal reasoning framework that separates image-only anatomical evidence acquisition from final clinical adjudication. Each frontal facial image is decomposed into seven fixed anatomical queries covering the neck, chin, mouth, face/neck fat, lower jaw, midface, and nose. Visual responses are converted into structured evidence cards recording target anatomy, visibility, risk direction, evidence strength, confidence, and a concise summary. These cards are combined with a cleaned clinical profile only in the final stage, where a large language model performs balanced binary screening adjudication. We evaluated EviOSAHS on a 642-subject cohort, mapping normal subjects to screening-negative and mild, moderate, or severe OSAHS subjects to screening-positive. EviOSAHS achieved 88.47% accuracy, 94.86% sensitivity, 93.74% F1-score, and a 5.14% false-negative rate, outperforming clinical-only prompting, direct multimodal prompting, and naive two-stage pipelines under a unified protocol. Ablations showed that seven-question visual decomposition and balanced final adjudication were critical to the high-sensitivity operating point. A question-level audit of 4,494 visual outputs showed a 100% structured parse rate and 93.88% high-visibility rate. EviOSAHS provides an auditable, high-sensitivity workflow for binary pre-polysomnography OSAHS screening, but should be viewed as a triage assistant rather than a diagnostic system. Prospective validation, external testing, and calibrated operating-point control are needed before clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。