针对文本流数据标签稀缺,提出新模型实现高效聚类与多标签学习。
Evolving Text Data Stream Mining
- 基于动态子空间学习,缓解高维文本数据对性能的压制。
- 在仅部分标注数据下实现主题演化捕捉与文档语义表征。
- 适合在线社交平台等实时文本分析场景,尤其标签难获取时。
文本流是随时间生成的有序文本文档序列,每日由在线社交平台产生海量此类数据。由于流数据具有无限长度、数据稀疏性和动态演化等特性,从这类数据中提取有用信息极具挑战性,尤其是在时间和内存受限的情况下。尽管过去十年提出了众多文本流挖掘算法,仍存在若干问题:首先,高维文本数据严重降低学习性能,现有方法需在子空间或缩减全局特征空间下运行;其次,如何提取文档的语义表示并捕捉随时间演化的主题仍具挑战;此外,标签稀缺问题普遍存在,而现有方法通常假设标签完全可用。为此,本文提出新的学习模型,用于文本流上的聚类和多标签学习,以应对标签稀缺问题。
原文摘要 · Abstract (English)
A text stream is an ordered sequence of text documents generated over time. A massive amount of such text data is generated by online social platforms every day. Designing an algorithm for such text streams to extract useful information is a challenging task due to unique properties of the stream such as infinite length, data sparsity, and evolution. Thereby, learning useful information from such streaming data under the constraint of limited time and memory has gained increasing attention. During the past decade, although many text stream mining algorithms have proposed, there still exists some potential issues. First, high-dimensional text data heavily degrades the learning performance until the model either works on subspace or reduces the global feature space. The second issue is to extract semantic text representation of documents and capture evolving topics over time. Moreover, the problem of label scarcity exists, whereas existing approaches work on the full availability of labeled data. To deal with these issues, in this thesis, new learning models are proposed for clustering and multi-label learning on text streams.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。