融合症状、咳嗽声、语音和胸片,构建轻量级肺炎筛查框架。
MultiSense-Pneumo: A Multimodal Learning Framework for Pneumonia Screening in Resource-Constrained Settings

- 整合多模态数据生成统一风险评估信号
- 影像路径在模拟域偏移下表现优异,但声音识别召回率低
- 适合资源受限地区研究使用,非临床诊断系统
肺炎仍是全球导致发病率和死亡率的主要疾病,尤其在缺乏影像设备、实验室检测和专科医疗资源的低资源地区。临床筛查依赖多种异构证据,包括症状、呼吸模式、语音描述和胸部影像,具有天然多模态特性。然而,现有计算方法多为单模态,主要聚焦于胸片分析。本文提出 MultiSense-Pneumo,一个面向肺炎筛查与分诊支持的多模态研究原型,融合结构化症状描述、咳嗽音频、语音内容及胸部X光片。系统采用确定性症状分诊、基于 LightGBM 的声学分类、使用 ResNet-18 的领域对抗影像分析、基于 Transformer 的语音识别,以及可解释的后融合算子。各模态转换为归一化关切信号后融合成统一筛查评分,融合权重手动设定,作为可解释的启发式参数而非学习或临床优化值。该系统设计为在标准笔记本硬件上离线运行,但未经过部署或临床验证。实验显示影像路径在合成域偏移下表现良好,但咳嗽声学异常类召回率下降,且缺乏端到端多模态患者评估。因此,MultiSense-Pneumo 定位为筛查与分诊研究的框架与组件级原型。
原文摘要 · Abstract (English)
Pneumonia remains a leading global cause of morbidity and mortality, particularly in low-resource settings where access to imaging, laboratory testing, and specialist care is limited. Clinical assessment relies on heterogeneous evidence, including symptoms, respiratory patterns, spoken descriptions, and chest imaging, making frontline screening inherently multimodal. However, many existing computational approaches remain unimodal and focus primarily on radiographs. In this work, we present MultiSense-Pneumo, a multimodal research prototype for pneumonia-oriented screening and triage support that integrates structured symptom descriptors, cough audio, spoken language, and chest radiographs. The system combines deterministic symptom triage, LightGBM-based acoustic classification, domain-adversarial radiograph analysis using ResNet-18, transformer-based speech recognition, and an interpretable late-fusion operator. Each modality is transformed into a normalized concern signal and aggregated into a unified screening estimate. The fusion weights are hand-specified and are treated as heuristic, interpretable parameters rather than learned or clinically optimized values. MultiSense-Pneumo is implemented with offline execution in mind on standard laptop-class hardware, but it is not presented as a deployment-validated or clinically validated diagnostic system. Experimental results demonstrate strong component-level performance of the radiograph pathway under synthetic domain shifts, while also highlighting important limitations, especially reduced abnormal-class recall for cough acoustics and the absence of paired end-to-end multimodal patient evaluation. MultiSense-Pneumo is therefore intended as a framework and component-level prototype for screening and triage research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。