用强化学习直接从对话生成医疗记录,准确率与可靠性双高。
Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge
- 用LoRA微调+DAPO强化学习,直接从语音生成SOAP病历
- 轻量/重模型均夺冠,概念匹配F1达最优,幻觉率最低
- 文本微调可迁移到语音,提升真实场景鲁棒性
本文介绍塔尔图大学在「超越转录挑战」(BeTraC)中的参赛系统,要求直接从长段医生-患者对话录音生成SOAP病历,无需中间转录。我们筛选了开源语音大模型的长音频鲁棒性,对Voxtral Mini(轻量级)和Voxtral Small(重量级)分别采用LoRA监督微调,再通过基于挑战指标(开放医学概念F1)的DAPO强化学习优化。两个系统均排名第一,独立评估显示其幻觉率低于所有参赛方案,表明以概念匹配为奖励的强化学习不损害事实准确性。此外发现,基于文本转录的微调能良好迁移到语音输入,显著提升对真实域外录音的鲁棒性。
原文摘要 · Abstract (English)
This paper describes TalTech's submissions to the Beyond Transcription Challenge (BeTraC), which requires generating SOAP notes directly from long doctor-patient conversation recordings, without intermediate transcription. After screening open-weight speech LLMs for long-audio robustness, we adapted Voxtral Mini (lightweight track) and Voxtral Small (heavyweight track) with LoRA supervised fine-tuning followed by DAPO reinforcement learning that uses the challenge metric, Open Medical Concept F1, as its reward. Our systems ranked first in both tracks, and an independent LLM-as-a-judge evaluation showed the lowest hallucination rate among all submissions, indicating that reinforcement learning against a concept-matching metric need not compromise factual reliability. We also find that fine-tuning on text transcripts transfers well to speech input and appears to improve robustness on out-of-domain real recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。