构建首个面向非正式口语的标点恢复数据集,提升真实场景下模型表现。
Spontaneous Informal Speech Dataset for Punctuation Restoration
- 从非正式口语中构建含标点和大小写的数据集
- 设计音视频与文本质量双重过滤机制,保障数据可信度
- 提供挑战性测试集,检验模型利用语音信息推断标点的能力
当前标点恢复模型主要在结构良好、剧本化语料上评估,但真实语音识别系统处理的多为包含大量不规则、结巴和语法偏差的非正式口语。为弥合这一差距,本文提出SponSpeech数据集,源自非正式口语来源,包含标点和大小写信息。除公开发布数据集外,还提供一个过滤管道,可评估语音音频与转录文本质量,用于生成更多高质量数据。同时,精心构建了一个“挑战性”测试集,旨在评估模型利用语音信息推断语法模糊标点的能力。SponSpeech已开源,代码及数据可通过 https://github.com/GitHubAccountAnonymous/PR 获取。
原文摘要 · Abstract (English)
Presently, punctuation restoration models are evaluated almost solely on well-structured, scripted corpora. On the other hand, real-world ASR systems and post-processing pipelines typically apply towards spontaneous speech with significant irregularities, stutters, and deviations from perfect grammar. To address this discrepancy, we introduce SponSpeech, a punctuation restoration dataset derived from informal speech sources, which includes punctuation and casing information. In addition to publicly releasing the dataset, we contribute a filtering pipeline that can be used to generate more data. Our filtering pipeline examines the quality of both speech audio and transcription text. We also carefully construct a ``challenging" test set, aimed at evaluating models' ability to leverage audio information to predict otherwise grammatically ambiguous punctuation. SponSpeech is available at https://github.com/GitHubAccountAnonymous/PR, along with all code for dataset building and model runs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。