arXiv:2602.12911cs.CL2026-02中稿 · LREC 2026被引 2

构建首个越语医学混用语音数据集,解决医疗场景中英词识别难题。

ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset & Benchmark

  • 构建34小时含1.6万条混用语句的越语医学语音数据集
  • 多语言预训练模型更擅长识别插入的英文医学术语
  • 结合越语优化与多语言预训练可平衡整体与混用识别准确率

在越南医疗沟通中,使用英语药物名或操作术语的语码转换现象普遍存在,给自动语音识别(ASR)系统带来挑战,尤其对低资源语言如越南语而言。现有大多数ASR系统难以正确识别越南语句子中的英文医学术语,且缺乏针对性基准。本文构建了一个34小时的越南医学语码转换语音数据集ViMedCSS,包含16,576个语句,每个语句至少包含一个来自五大学科领域双语词典的英文医学术语。基于该数据集,我们评估了多个前沿ASR模型,并测试不同微调策略以提升医学术语识别效果。实验表明,越南语优化模型在通用段落上表现更好,而多语言预训练有助于捕捉英文插入;两者结合能实现整体与混用识别准确率的最佳平衡。本工作首次提供越语医学语码转换的基准,为低资源多语言ASR的领域适配提供了有效思路。

原文摘要 · Abstract (English)

Code-switching (CS), which is when Vietnamese speech uses English words like drug names or procedures, is a common phenomenon in Vietnamese medical communication. This creates challenges for Automatic Speech Recognition (ASR) systems, especially in low-resource languages like Vietnamese. Current most ASR systems struggle to recognize correctly English medical terms within Vietnamese sentences, and no benchmark addresses this challenge. In this paper, we construct a 34-hour Vietnamese Medical Code-Switching Speech dataset (ViMedCSS) containing 16,576 utterances. Each utterance includes at least one English medical term drawn from a curated bilingual lexicon covering five medical topics. Using this dataset, we evaluate several state-of-the-art ASR models and examine different specific fine-tuning strategies for improving medical term recognition to investigate the best approach to solve in the dataset. Experimental results show that Vietnamese-optimized models perform better on general segments, while multilingual pretraining helps capture English insertions. The combination of both approaches yields the best balance between overall and code-switched accuracy. This work provides the first benchmark for Vietnamese medical code-switching and offers insights into effective domain adaptation for low-resource, multilingual ASR systems.

语音识别医学AI语码转换低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。