arXiv:2608.14150cs.CL2026-08

通过静音裁剪与合成监督提升多语种对话理解性能

Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge

论文配图:Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge
图 1 · 摘自论文原文
  • 用随机前导静音裁剪和动态平均策略优化语音识别
  • 合成问答对并结合标签直接回答,使准确率提升至86.0%
  • 适合关注多语种对话系统、语音增强的工程师

第二届多语言对话语音语言模型(MLC-SLM)挑战赛评估两个任务:说话人分离与识别(任务1)以及对话语音理解(任务2)。两个任务均在无原始语音段边界和说话人标签的情况下进行评估,任务2亦无问答训练集。针对任务1,我们使用VibeVoice-ASR-7B模型,采用随机前导静音裁剪、一致的时间戳修正及指数移动平均(EMA)训练策略。针对任务2,通过多模态候选生成、静音音频过滤与分布匹配增强构建合成问答对,并对Qwen3-Omni-30B-A3B-Instruct模型进行带标签直接回答微调。在任务1测试集上,裁剪使tcpMER从18.30%降至17.27%,EMA进一步降至16.73%。在任务2测试集上,联合使用分布匹配增强与标签直接回答,准确率从83.0%提升至86.0%。

原文摘要 · Abstract (English)

The second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge evaluates two tasks over complete, unsegmented multilingual conversations: speaker diarization and recognition (Task 1) and conversational speech understanding (Task 2). Neither task provides oracle utterance boundaries or speaker labels at evaluation, and Task 2 provides no question-answer training set. For Task 1, we fine-tune VibeVoice-ASR-7B with random leading-silence cropping, consistent timestamp correction, and an exponential moving average (EMA) training strategy. For Task 2, we construct synthetic question-answer pairs through multimodal candidate generation, silent-audio filtering, and distribution-matched augmentation, and fine-tune Qwen3-Omni-30B-A3B-Instruct for tagged direct answering. On the Task 1 evaluation set, cropping reduces tcpMER from 18.30% to 17.27%, and EMA further reduces it to 16.73%. On the Task 2 evaluation set, jointly applying distribution-matched augmentation and tagged direct answering raises accuracy from 83.0% to 86.0%.

多语言对话语音识别合成数据模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。