用少量标注数据实现医疗音频的智能诊断,无需大量训练数据。
Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio

- 通过聚类生成伪标签,实现无监督上下文学习
- 在2路2样本测试中达71.6%准确率,比基线高9%以上
- 适合资源匮乏地区的医院部署,支持跨机构协作
低资源环境下临床音频诊断需要仅依赖少量示例即可识别疾病的模型。我们提出联邦自上下文化(FSC)框架,一种用于跨联邦医院客户端进行上下文临床音频诊断的多模态语言模型。FSC通过音频表征的无监督聚类构建伪标签样本集,绕过稀缺的真实诊断标签,并实现从支持-查询对中的上下文推理。其渐进式三阶段流程首先通过基于字幕的预训练对齐音频嵌入与语言模型,再通过联邦优化适配其周期性上下文推理能力。测试时,给定少量标注的支持集,模型通过多模态推理诊断未见查询。在保留的呼吸与心脏疾病上,FSC在2路2样本评估中达到71.6%准确率,优于音频-语言基线超过9%。
原文摘要 · Abstract (English)
Clinical audio diagnosis in low-resource settings requires models that identify conditions from minimal examples without large annotated corpora. We propose Federated Self-Contextualization (FSC), a multimodal language model framework for in-context clinical audio diagnosis across federated hospital clients. FSC constructs pseudo-label episodes via unsupervised clustering of audio representations, bypassing scarce real diagnostic labels, and enables contextual reasoning from support-query pairs. Our progressive three-stage pipeline first aligns audio embeddings with the language model via caption-based pretraining, then adapts it for episodic in-context inference through federated optimization. At test time, given a small labeled support set, the model diagnoses an unseen query through multimodal reasoning. On held-out respiratory and cardiac conditions, FSC achieves 71.6% accuracy in 2-way 2-shot evaluation, outperforming audio-language baselines by over 9%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。