arXiv:2601.02128cs.CLeess.AS2026-01被引 3

用LoRA微调大模型,实现会议与讲座文本的多层级目录自动生成。

Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation

  • 基于LoRA微调大模型,结合语音停顿特征进行分层话题分割
  • 在英、葡、德语会议和讲座数据上显著超越现有基线
  • 提出统一评估多层级结构的新指标,适合内容组织与无障碍应用

将语音转录文本分割为主题段落,有助于下游处理及依赖文字的可访问性需求。本文提出一种新型层次化话题分割方法,生成包含主题与子主题边界的多层级目录。对比了零样本提示与LoRA微调在大语言模型上的效果,并探索了高阶语音停顿特征的融合。在英语会议录音与多语言讲座转录(葡萄牙语、德语)上的实验表明,该方法显著优于现有话题分割基线。此外,我们改进了一种通用评估指标,使其能够在一个度量中同时考虑所有层级的分割效果。

原文摘要 · Abstract (English)

Segmenting speech transcripts into thematic sections benefits both downstream processing and users who depend on written text for accessibility. We introduce a novel approach to hierarchical topic segmentation in transcripts, generating multi-level tables of contents that capture both topic and subtopic boundaries. We compare zero-shot prompting and LoRA fine-tuning on large language models, while also exploring the integration of high-level speech pause features. Evaluations on English meeting recordings and multilingual lecture transcripts (Portuguese, German) show significant improvements over established topic segmentation baselines. Additionally, we adapt a common evaluation measure for multi-level segmentation, taking into account all hierarchical levels within one metric.

文本分割LoRA多层级语音转录

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。