用少量合成语音数据让语音模型高效完成医疗问诊对话。
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation
- 分两阶段训练:先用文本注入医学知识,再用少量语音数据对齐模态。
- 仅需10k合成语音样本,即可在多轮问诊中超越所有基线模型。
- 适合需要轻量化部署医疗语音助手的研究与开发者。
医疗问诊本质上以语音为核心,但以往研究多聚焦长文本交互,繁琐且不友好。近期语音语言模型(SpeechLMs)实现了更自然的语音交互,然而医学语音数据稀缺以及直接微调效率低下,制约了其在医疗问诊中的应用。本文提出 SpeechMedAssist,一种原生支持语音多轮交互的 SpeechLM。通过利用 SpeechLM 架构特性,将传统单阶段训练分解为两阶段:(1) 通过文本注入知识与能力,(2) 仅用10k合成语音数据进行模态重对齐,大幅降低对真实医疗语音数据的需求。我们设计了一个包含单轮问答与多轮模拟交互的基准测试。实验表明,该模型在多数评估场景中均优于所有基线,在有效性和鲁棒性上表现更优。
原文摘要 · Abstract (English)
Medical consultations are intrinsically speech-centric. However, most prior works focus on long-text-based interactions, which are cumbersome and patient-unfriendly. Recent advances in speech language models (SpeechLMs) have enabled more natural speech-based interaction, yet the scarcity of medical speech data and the inefficiency of directly fine-tuning on speech data jointly hinder the adoption of SpeechLMs in medical consultation. In this paper, we propose SpeechMedAssist, a SpeechLM natively capable of conducting speech-based multi-turn interactions with patients. By exploiting the architectural properties of SpeechLMs, we decouple the conventional one-stage training into a two-stage paradigm consisting of (1) Knowledge & Capability Injection via Text and (2) Modality Re-alignment with Limited Speech Data, thereby reducing the requirement for medical speech data to only 10k synthesized samples. To evaluate SpeechLMs for medical consultation scenarios, we design a benchmark comprising both single-turn question answering and multi-turn simulated interactions. Experimental results show that our model outperforms all baselines in both effectiveness and robustness in most evaluation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。