提出适用于依赖文本的变点检测方法,理论+实证双验证。
Consistent Kernel Change-Point Detection under m-Dependence for Text Segmentation
- 在m-依赖条件下证明变点数量一致性与位置弱一致性
- 使用LLM生成合成文本验证渐近性质,效果优于基线
- 首个基于现代嵌入的文本分割实证研究,适合自然语言分析者
核变点检测(KCPD)已成为识别复杂数据结构变化的常用工具。尽管现有理论在独立性假设下建立了相合性,但真实序列数据(如文本)具有强依赖性。本文在m-依赖条件下建立新理论保证:在较弱额外假设下,证明了变点数量的一致性及位置的弱一致性。通过基于大模型(LLM)生成的合成m-依赖文本进行模拟验证渐近性质。为进一步补充,首次系统地对现代嵌入下的文本分割任务开展全面实证研究。在多个文本数据集上,基于文本嵌入的KCPD在标准分割指标上表现优于基线。案例研究以泰勒·斯威夫特推文为例,表明KCPD兼具理论可靠性与实际有效性。
原文摘要 · Abstract (English)
Kernel change-point detection (KCPD) has become a widely used tool for identifying structural changes in complex data. While existing theory establishes consistency under independence assumptions, real-world sequential data such as text exhibits strong dependencies. We establish new guarantees for KCPD under $m$-dependent data: specifically, we prove consistency in the number of detected change points and weak consistency in their locations under mild additional assumptions. We perform an LLM-based simulation that generates synthetic $m$-dependent text to validate the asymptotics. To complement these results, we present the first comprehensive empirical study of KCPD for text segmentation with modern embeddings. Across diverse text datasets, KCPD with text embeddings outperforms baselines in standard text segmentation metrics. We demonstrate through a case study on Taylor Swift's tweets that KCPD not only provides strong theoretical and simulated reliability but also practical effectiveness for text segmentation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。