用LoRA微调大模型,实现会议与讲座文本的多层级目录自动生成。
Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation
- 基于LoRA微调大模型,结合语音停顿特征进行分层话题分割
- 在英、葡、德语会议和讲座数据上显著超越现有基线
- 提出统一评估多层级结构的新指标,适合内容组织与无障碍应用
将语音转录文本分割为主题段落,有助于下游处理及依赖文字的可访问性需求。本文提出一种新型层次化话题分割方法,生成包含主题与子主题边界的多层级目录。对比了零样本提示与LoRA微调在大语言模型上的效果,并探索了高阶语音停顿特征的融合。在英语会议录音与多语言讲座转录(葡萄牙语、德语)上的实验表明,该方法显著优于现有话题分割基线。此外,我们改进了一种通用评估指标,使其能够在一个度量中同时考虑所有层级的分割效果。
原文摘要 · Abstract (English)
Segmenting speech transcripts into thematic sections benefits both downstream processing and users who depend on written text for accessibility. We introduce a novel approach to hierarchical topic segmentation in transcripts, generating multi-level tables of contents that capture both topic and subtopic boundaries. We compare zero-shot prompting and LoRA fine-tuning on large language models, while also exploring the integration of high-level speech pause features. Evaluations on English meeting recordings and multilingual lecture transcripts (Portuguese, German) show significant improvements over established topic segmentation baselines. Additionally, we adapt a common evaluation measure for multi-level segmentation, taking into account all hierarchical levels within one metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。