arXiv:2512.24517cs.CL2025-12中稿 · LREC 2026被引 1

为语音转写文本添加段落分隔,提升可读性和复用性

Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech

  • 提出约束解码法,让大模型在不改动原文的前提下插入段落断点
  • 构建了两个新基准数据集,分别来自真实演讲和合成标签视频
  • 小型模型实现最佳效果,还能低成本同时预测章节与段落

自动语音转写通常以无结构的词流形式输出,影响可读性和再利用。本文将段落分割视为缺失的结构化步骤,填补了语音处理与文本分割交叉领域的三个空白:首先,构建了首个针对该任务的基准数据集——人工标注的TEDPara(TED演讲)和合成标签的YTSegPara(YouTube视频),聚焦未被充分研究的语音领域;其次,提出一种约束解码方法,使大语言模型能在保持原转录内容完整性的前提下插入段落分隔符,支持精准、逐句评估;第三,展示小型模型MiniSeg达到当前最优性能,并可通过层级扩展,以极低计算成本联合预测章节与段落。整体工作确立了段落分割在语音处理中的标准化地位。

原文摘要 · Abstract (English)

Automatic speech transcripts are often delivered as unstructured word streams that impede readability and repurposing. We recast paragraph segmentation as the missing structuring step and fill three gaps at the intersection of speech processing and text segmentation. First, we establish TEDPara (human-annotated TED talks) and YTSegPara (YouTube videos with synthetic labels) as the first benchmarks for the paragraph segmentation task. The benchmarks focus on the underexplored speech domain, where paragraph segmentation has traditionally not been part of post-processing, while also contributing to the wider text segmentation field, which still lacks robust and naturalistic benchmarks. Second, we propose a constrained-decoding formulation that lets large language models insert paragraph breaks while preserving the original transcript, enabling faithful, sentence-aligned evaluation. Third, we show that a compact model (MiniSeg) attains state-of-the-art accuracy and, when extended hierarchically, jointly predicts chapters and paragraphs with minimal computational cost. Together, our resources and methods establish paragraph segmentation as a standardized, practical task in speech processing.

语音处理文本分割大模型应用数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。