用大模型生成能力提升文本主题切分效果
Topic Segmentation Using Generative Language Models
- 采用递归枚举提示策略,捕捉长距离依赖关系
- 在边界相似性指标下表现优于传统方法
- 适合需要高精度主题划分的学术与内容分析场景
使用生成式大语言模型进行主题分割仍处于探索阶段。以往方法依赖句子间的语义相似性,但此类模型缺乏大模型所具备的长程依赖建模能力和广泛知识。本文提出一种基于句子枚举的重叠递归提示策略,并支持边界相似性评估指标。实验结果表明,大语言模型在主题分割任务上表现优于现有方法,但仍存在若干待解决的问题,尚不能完全依赖其进行主题分割。
原文摘要 · Abstract (English)
Topic segmentation using generative Large Language Models (LLMs) remains relatively unexplored. Previous methods use semantic similarity between sentences, but such models lack the long range dependencies and vast knowledge found in LLMs. In this work, we propose an overlapping and recursive prompting strategy using sentence enumeration. We also support the adoption of the boundary similarity evaluation metric. Results show that LLMs can be more effective segmenters than existing methods, but issues remain to be solved before they can be relied upon for topic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。