arXiv:2506.18510cs.SDcs.AI2025-06中稿 · INTERSPEECH2025 wo…被引 3

用不完美的语音提示,让大模型生成带停顿标注的完整转录。

Smooth Operators: LLMs Translating Imperfect Hints into Disfluency-Rich Transcripts

  • 用音频编码器+文本输入,让大模型把有缺陷的转录变成带停顿标注的完整文本。
  • 即使输入文本有错误,只要带时间戳,模型也能生成准确的停顿标注结果。
  • 适合做语音识别、口语分析或无障碍技术的研究者使用。

在语音与语言处理系统中,准确检测口语中的停顿对提升性能和推动包容性技术发展至关重要。我们提出一种新方法,利用大语言模型(LLMs)作为多功能学习者,将非词汇性输入(如音频、视频)与文本输入结合,将停顿以显式标记形式加入转录,并附带时间戳,生成完整标注的停顿丰富型转录。该方法融合音频编码器提取的声学特征,以及不同质量的文本输入:无停顿的干净转录、对齐工具生成的时间对齐转录,或基于音素的自动语音识别(ASR)模型输出——这些输入均可能包含缺陷。关键发现是,文本输入无需完美,只要包含时间戳信息,大模型就能有效修正并生成完整的停顿标注转录,体现出其对不完美提示的强大鲁棒性。

原文摘要 · Abstract (English)

Accurate detection of disfluencies in spoken language is crucial for enhancing the performance of automatic speech and language processing systems, as well as fostering the development of more inclusive speech and language technologies. Leveraging the growing trend of large language models (LLMs) as versatile learners capable of processing both lexical and non-lexical inputs (e.g., audio and video), we propose a novel approach to transcribing disfluencies as explicit tokens with timestamps, enabling the generation of fully annotated disfluency-rich transcripts. Our method integrates acoustic representations extracted from an audio encoder with textual inputs of varying quality: clean transcriptions without disfluencies, time-aligned transcriptions from aligners, or outputs from phoneme-based ASR models -- all of which may contain imperfections. Importantly, our experiments demonstrate that textual inputs do not need to be flawless. As long as they include timestamp-related cues, LLMs can effectively smooth the input and produce fully disfluency-annotated transcripts, underscoring their robustness in handling imperfect hints.

语音识别大模型停顿检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。