arXiv:2411.00876cs.LGcs.AI2024-11被引 2

提出融合分类与聚类的开放集识别框架,应对数据流中未知类别的挑战。

Resilience to the Flowing Unknown: an Open Set Recognition Framework for Data Streams

  • 结合分类与聚类,动态识别数据流中的未知类别。
  • 在多数据集基准上优于传统增量分类器,尤其在未知类比例高时表现更稳。
  • 适合需要长期运行、面对未知场景的智能系统开发者参考。

现代数字应用广泛集成人工智能模型以实现自动化决策,但在复杂动态场景下持续生成的数据流中,这些AI系统面临可靠性和安全性挑战。本文研究韧性AI系统,需应对训练中未见的意外情况。传统闭集分类器在流式场景中存在‘过占空间’问题,即强制将所有新样本归入已知类别。开放集识别研究在批量学习中已针对此问题提出解决方案。本文提出一种结合分类与聚类的开放集识别框架,用于解决流式环境下的该难题。构建了包含不同已知/未知类别比例的数据集基准,系统性对比所提混合框架与单一增量分类器的性能。实验结果表明,所提框架在未知类比例较高时表现更优;同时揭示了增量分类器在开放世界流式环境中面临的局限与障碍。

原文摘要 · Abstract (English)

Modern digital applications extensively integrate Artificial Intelligence models into their core systems, offering significant advantages for automated decision-making. However, these AI-based systems encounter reliability and safety challenges when handling continuously generated data streams in complex and dynamic scenarios. This work explores the concept of resilient AI systems, which must operate in the face of unexpected events, including instances that belong to patterns that have not been seen during the training process. This is an issue that regular closed-set classifiers commonly encounter in streaming scenarios, as they are designed to compulsory classify any new observation into one of the training patterns (i.e., the so-called \textit{over-occupied space} problem). In batch learning, the Open Set Recognition research area has consistently confronted this issue by requiring models to robustly uphold their classification performance when processing query instances from unknown patterns. In this context, this work investigates the application of an Open Set Recognition framework that combines classification and clustering to address the \textit{over-occupied space} problem in streaming scenarios. Specifically, we systematically devise a benchmark comprising different classification datasets with varying ratios of known to unknown classes. Experiments are presented on this benchmark to compare the performance of the proposed hybrid framework with that of individual incremental classifiers. Discussions held over the obtained results highlight situations where the proposed framework performs best, and delineate the limitations and hurdles encountered by incremental classifiers in effectively resolving the challenges posed by open-world streaming environments.

开放集识别数据流增量学习聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。