用合成语音和自动生成的SOAP标签,实现语音直接转医疗摘要。
Speech-to-SOAP: End-to-End Summarization of Medical Dialogues: KIT@BeTraC 2026
- 通过合成语音和自动标注构建统一数据集,提升模型适应性。
- 端到端语音转SOAP摘要,比传统转录方式更快更完整。
- 适合医疗自动化、语音处理研究者,尤其关注临床记录效率提升。
随着大语言模型指令遵循能力的发展,摘要生成成为热门应用。其中,临床病历提取任务尤为受关注,可显著减少医护人员的文书负担,使其专注于核心医疗工作。进一步自动化方向是直接从语音生成临床记录,跳过中间文本转录环节,从而保留咳嗽等副语言特征,避免信息丢失。为此,我们提交了KIT在2026年BeTraC挑战赛轻量级赛道的方案。主要贡献是一个可扩展的数据增强管道,通过合成语音与自动生成的SOAP监督信号,统一异构的医学对话数据集,使语音基础模型能稳健地实现端到端语音转SOAP摘要。
原文摘要 · Abstract (English)
With the advent of Large Language Models and its instruction following capabilities a promising application is the task of summarization. Within this domain of task the extractive sub-task of clinical protocolling has emerged as a topic of particular interest as it can significantly reduce the downtime and protocolling burden of health-care workers thus enabling them to focus on their core work helping humans. A further step towards automation is the direct generation of clinical notes from speech without intermediate transcripts, reducing processing time while preserving information such as coughing or other paralinguistic cues that may be lost in transcript-based systems. To this end, we present KIT's submission to this years BeTraC challenge in the lightweight track. Our main contribution is a scalable data augmentation pipeline that unifies heterogeneous medical dialogue datasets through synthetic speech generation and automatically generated SOAP supervision, enabling robust adaptation of a speech foundation model for end-to-end speech-to-SOAP generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。