arXiv:2511.01619cs.CL2025-11中稿 · LREC 2026 conferen…被引 2

扩充4种斯拉夫语议会语音语料库,自动添加多层标注。

ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian

  • 基于转录文本与录音对齐,自动添加语言学与情感标注。
  • 四语语料均含填充停顿标记,两语额外提供音节与重音标注。
  • 适用于语音分析、情感识别及跨语言研究的高质量数据集。

ParlaSpeech 是涵盖克罗地亚语、捷克语、波兰语和塞尔维亚语四种斯拉夫语的议会语音语料库,总计6000小时。本版本通过自动对齐帕尔拉明特(ParlaMint)转录文本与对应议会录音,大幅丰富了各语料库的标注层。所有语料在文本层面增加了语言学标注与情感预测;语音层面自动标注了最常见的言语不流畅现象——填充停顿。其中两种语言还额外提供了词级与音素级对齐,以及多音节词主重音位置的自动标注。这些增强显著提升了语料在多个研究领域的应用价值,本文通过声学特征与情感关联的分析予以验证。所有数据以JSONL和TextGrid格式公开,并可通过语料检索工具访问。

原文摘要 · Abstract (English)

ParlaSpeech is a collection of spoken parliamentary corpora currently spanning four Slavic languages - Croatian, Czech, Polish and Serbian - all together 6 thousand hours in size. The corpora were built in an automatic fashion from the ParlaMint transcripts and their corresponding metadata, which were aligned to the speech recordings of each corresponding parliament. In this release of the dataset, each of the corpora is significantly enriched with various automatic annotation layers. The textual modality of all four corpora has been enriched with linguistic annotations and sentiment predictions. Similar to that, their spoken modality has been automatically enriched with occurrences of filled pauses, the most frequent disfluency in typical speech. Two out of the four languages have been additionally enriched with detailed word- and grapheme-level alignments, and the automatic annotation of the position of primary stress in multisyllabic words. With these enrichments, the usefulness of the underlying corpora has been drastically increased for downstream research across multiple disciplines, which we showcase through an analysis of acoustic correlates of sentiment. All the corpora are made available for download in JSONL and TextGrid formats, as well as for search through a concordancer.

语音语料多语言自动标注情感分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。