arXiv:2501.07723cs.CLcs.LG2025-01被引 1

用词法与字符ngram特征,简单高效分割文本基本语篇单元。

ESURF: Simple and Effective EDU Segmentation

  • 基于词法和字符n-gram特征,用随机森林判断语篇边界。
  • 在多个数据集上优于现有方法,且提升主流语篇分析器性能。
  • 方法简洁有效,适合追求训练效率的语篇分析任务。

将文本分割为基本语篇单元(EDUs)是语篇分析的基础任务。本文提出一种基于词法与字符n-gram特征、采用随机森林分类的简单方法,用于识别EDU边界并完成分割。尽管方法简单,其在分割任务及作为状态领先语篇解析器组件时均表现更优,表明这些特征对识别基本语篇单元至关重要,提示未来可能发展出更训练高效的语篇分析方法。

原文摘要 · Abstract (English)

Segmenting text into Elemental Discourse Units (EDUs) is a fundamental task in discourse parsing. We present a new simple method for identifying EDU boundaries, and hence segmenting them, based on lexical and character n-gram features, using random forest classification. We show that the method, despite its simplicity, outperforms other methods both for segmentation and within a state of the art discourse parser. This indicates the importance of such features for identifying basic discourse elements, pointing towards potentially more training-efficient methods for discourse analysis.

语篇分析文本分割随机森林

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。