arXiv:2410.05361cs.LGcs.AI2024-10NeurIPS被引 22

用大模型统一分析呼吸音和文本,提升疾病预测准确率

RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction

  • 融合音频与文本的跨模态注意力机制
  • 在五个数据集上平均提升4.6%准确率,新任务零样本预测成功
  • 适合医疗健康、多模态诊断领域研究者

呼吸系统疾病发病率和死亡率高,早期筛查至关重要。机器学习可辅助临床问诊与听诊,但涉及人口统计、病史、症状及呼吸音频等异构复杂数据,现有方法受限于训练数据少、融合方式简单、任务专用模型,泛化能力不足。本文提出RespLLM,一种新型多模态大语言模型框架,通过交叉模态注意力统一文本与音频表征,利用预训练LLM的先验知识实现高效融合。通过指令微调整合多源数据,确保模型通用性与灵活性。在五个真实世界数据集上的实验表明,RespLLM在训练任务上平均优于基线4.6%,在未见数据集上提升7.9%,并可实现新任务的零样本预测。本工作为能感知、聆听、理解异构数据的多模态模型奠定基础,推动可扩展的呼吸健康诊断发展。

原文摘要 · Abstract (English)

The high incidence and mortality rates associated with respiratory diseases underscores the importance of early screening. Machine learning models can automate clinical consultations and auscultation, offering vital support in this area. However, the data involved, spanning demographics, medical history, symptoms, and respiratory audio, are heterogeneous and complex. Existing approaches are insufficient and lack generalizability, as they typically rely on limited training data, basic fusion techniques, and task-specific models. In this paper, we propose RespLLM, a novel multimodal large language model (LLM) framework that unifies text and audio representations for respiratory health prediction. RespLLM leverages the extensive prior knowledge of pretrained LLMs and enables effective audio-text fusion through cross-modal attentions. Instruction tuning is employed to integrate diverse data from multiple sources, ensuring generalizability and versatility of the model. Experiments on five real-world datasets demonstrate that RespLLM outperforms leading baselines by an average of 4.6% on trained tasks, 7.9% on unseen datasets, and facilitates zero-shot predictions for new tasks. Our work lays the foundation for multimodal models that can perceive, listen to, and understand heterogeneous data, paving the way for scalable respiratory health diagnosis.

多模态呼吸健康大模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。