arXiv:2506.03793cs.CL2025-06被引 2

用大模型提升多语言语音转写标点恢复准确率

Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts

  • 基于预训练大模型构建通用标点恢复模型
  • 支持22种印度语言和英语,性能超越现有方法
  • 适合低资源场景的语音处理与下游任务应用

标点对语义结构至关重要,但现有模型在处理口语转录中的不流畅现象(如重复、回溯)时表现不佳,影响翻译、语音合成、摘要等下游任务质量。本文提出Cadence,一个基于预训练大语言模型的通用标点恢复模型,能同时处理干净书面语和高度自发的口语转录。其性能超越当前最佳水平,并将语言支持从14种扩展至22种印度语言及英语。我们对不同标点类型和语系下的模型行为进行了全面分析,发现领域迁移和稀有标点仍存在挑战。结果表明,利用预训练语言模型进行多语言标点恢复有效可行,且在大规模低资源NLP流程中具有实用价值。

原文摘要 · Abstract (English)

Punctuation plays a vital role in structuring meaning, yet current models often struggle to restore it accurately in transcripts of spontaneous speech, especially in the presence of disfluencies such as false starts and backtracking. These limitations hinder the performance of downstream tasks like translation, text to speech, summarization, etc. where sentence boundaries are critical for preserving quality. In this work, we introduce Cadence, a generalist punctuation restoration model adapted from a pretrained large language model. Cadence is designed to handle both clean written text and highly spontaneous spoken transcripts. It surpasses the previous state of the art in performance while expanding support from 14 to all 22 Indian languages and English. We conduct a comprehensive analysis of model behavior across punctuation types and language families, identifying persistent challenges under domain shift and with rare punctuation marks. Our findings demonstrate the efficacy of utilizing pretrained language models for multilingual punctuation restoration and highlight Cadence practical value for low resource NLP pipelines at scale.

标点恢复多语言语音转录大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。