让语音大模型直接预测字级时间戳,提升字幕与媒体同步精度。
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
- 在语音大模型中直接加入字级时间戳预测模块
- 新训练策略使时间戳准确率提升,同时改善整体识别效果
- 适合需要精准时间对齐的视频字幕、媒体搜索场景
近期语音感知语言模型将强大的声学编码器与大语言模型结合,使系统从单纯转录迈向生成更丰富输出。其中,字级时间戳预测对字幕生成、媒体搜索和多模态同步至关重要,但通常依赖外部对齐工具。本文扩展现有语音感知语言模型,使其直接联合预测文本与时间戳。提出一组轻量级新训练策略,在保持识别质量的同时增强对齐鲁棒性。跨多个数据集的实验表明,这些策略不仅提升时间戳准确性,还带来整体自动语音识别性能的提升。结果证明了一种高效统一的语音识别与精确时间戳预测方案。
原文摘要 · Abstract (English)
Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer outputs. Among these, word-level timestamp prediction is critical for applications such as captioning, media search, and multimodal synchronization, yet it is often handled by external alignment tools. In this work, we extend an existing speech-aware language model to predict timestamps directly alongside transcripts. We introduce a set of novel lightweight training strategies that improve alignment robustness while preserving recognition quality. Experiments across multiple datasets show that these strategies not only enhance timestamp accuracy, but also yield gains in overall ASR performance. Together, they demonstrate an efficient and unified approach to speech recognition with precise timestamp prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。