开源105万参数模型,专用于高精度医疗语音转写
MedASR: An Open-Source Model for High-Accuracy Medical Dictation
- 小模型设计,用伪流式滑窗提升长文本转录准确率
- 在眼动追踪数据上比Whisper Large-v3降低58%错误率
- 适合医疗领域开发者快速搭建私有化语音系统
我们提出MedASR,一个面向高精度医疗语音转写的开源105M参数模型。该模型遵循‘小、快、准’设计理念,聚焦三大核心问题:(1)临床语料稀缺与类别不平衡;(2)高效长文本训练;(3)通过伪流式滑窗机制实现高精度推理。评估显示,相较于Whisper Large-v3,MedASR在Eye Gaze数据集上实现了58%的相对词错误率(WER)降低。通过开源MedASR,我们为专业医疗应用提供透明且高性能的底层模型,打破专有系统对临床文档工作的壁垒。
原文摘要 · Abstract (English)
We present MedASR, an open-source 105M-parameter model engineered for high-accuracy medical dictation. Prioritizing a "small, fast, and accurate" design, MedASR addresses 3 core pillars (1) Data: overcoming clinical corpora scarcity and class imbalance; (2) Modeling: efficient long-form training; and (3) Inference: accurate transcription via a pseudo-streaming sliding-window approach. Our evaluation shows that MedASR achieves a 58% relative WER reduction on Eye Gaze compared to Whisper Large-v3. By open-sourcing MedASR, we provide a transparent, high-performance backbone for specialized health-care applications, breaking down the barriers to clinical documentation often obscured by proprietary systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。