arXiv:2605.16545cs.LGcs.AI2026-05

打造实时医疗语音识别系统,精准处理临床术语与缩写。

Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces

论文配图:Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces
图 1 · 摘自论文原文
  • 分模块设计:识别、格式化、上下文修正协同优化
  • 医疗数据集上性能显著优于现有系统,通用场景也表现优异
  • 适合临床医生实时记录、病历生成等安全敏感场景

尽管语音在医疗领域已用于转录和环境记录多年,但医学语音识别仍面临挑战:需准确捕捉专业术语、解析上下文歧义,并精确还原测量值、缩写与临床简写。现有方案多针对通用转录或特定文书流程优化,限制了其在高安全要求环境中的可靠性及更广泛临床应用的价值。本文提出 Symphony for Speech-to-Text,一个面向实时流式与批量文件处理的医疗级语音识别系统。该系统将转录过程分解为识别、格式化与上下文校正三个专用组件,在保证临床术语召回率的同时实现实时生成结构化文本,并适应不同使用场景。在公开基准与医疗语音数据集上的评估显示,Symphony 在临床场景中显著超越当前最优系统,且在通用领域表现持平或更优,表明其具备稳健泛化能力而非过拟合。研究团队发布了首个临床基准数据集,以支持可靠验证与领域进步。Symphony 已通过生产级 API 发布,支持实时口述、对话转录与批量音频处理。

原文摘要 · Abstract (English)

After decades of use in dictation and, more recently, ambient documentation, speech is emerging as a primary modality for interacting with technology and AI in healthcare. Yet medical speech recognition remains difficult: systems must capture specialized terminology, resolve contextual ambiguity, and render measurements, abbreviations, and clinical shorthand precisely. Existing solutions are typically optimized either for general-purpose transcription or narrow dictation workflows, limiting their reliability in safety-critical settings and their usefulness for broader clinical workflows. We introduce Symphony for Speech-to-Text, a medical-grade speech recognition system for real-time streaming and batch file-based clinical use. Symphony decomposes the transcription process into specialized components for recognition, formatting, and contextual correction to optimize medical term recall while producing clinically structured text in real time and adapting across use cases. Evaluations on public benchmark and medical speech datasets show that Symphony substantially outperforms state-of-the-art systems in clinical settings while matching or exceeding them in general-domain settings, suggesting robust generalization rather than overfitting. We release a clinical benchmark dataset to support reliable validation and further progress in medical speech recognition. Symphony is available through a production-grade API for live dictation, conversational transcription, and batch audio file processing.

语音识别医疗AI实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。