arXiv:2607.18934cs.CL2026-07中稿 · Interspeech 2026 l…

让语音识别模型按需输出原文或意译,解决风格混乱问题

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

  • 用双语对照数据训练解码器,让模型学会控制输出风格
  • 零样本提升德语不流畅检测准确率至79%,英语微调更优
  • 可生成高质量原文转录,适合构建高精度语音语料库

当前语音识别模型在异构标注数据上训练时,将转录风格(原文或意译)视为不可控的隐变量,导致解码不稳定、评估混淆(高达60%的错误率由风格不匹配引起)以及词级时间戳不可靠。我们发现模型已编码两种风格,关键在于可控激活。通过在并行原文/意译对上训练带覆盖感知的解码任务标记,模型在仅英语训练下实现德语不流畅检测F1从10%提升至79%(零样本)。全英文微调后,在原文准确性、不流畅检测和意译质量上均超越所有基线,跨语言表现优异。进一步引入监督交叉注意力微调,使不流畅语音的词级时间戳优于强制对齐基线。最后提出verbatimize新任务,支持高效构建与增强高质量原始转录语料库。

原文摘要 · Abstract (English)

Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER attributable to style mismatch), and unreliable word-level timing. We show that models already encode both styles; the challenge is controlled activation. Using coverage-aware decoder task tokens trained on parallel verbatim/intended transcript pairs, we raise German disfluency F1 from 10% to 79% zero-shot, despite English-only training. Full English-only fine-tuning surpasses all baselines in verbatim accuracy, disfluency detection, and intended-mode quality across both languages. We further introduce supervised cross-attention fine-tuning that improves word-level timestamps on disfluent speech beyond forced-alignment baselines. Finally, we propose verbatimize, a new task enabling scalable creation and enrichment of speech corpora with high-quality canonical verbatim transcriptions.

语音识别风格控制转录生成时间戳优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。