解决流数据中混合特征、分布漂移和标签缺失的在线学习难题
Extension OL-MDISF: Online Learning from Mix-Typed, Drifted, and Incomplete Streaming Features
- 用耦合模型构建统一潜在空间,适应多种特征类型
- 自适应滑动窗口检测概念漂移,保持模型稳定
- 基于几何关系推断标签近似,降低标注依赖
在线学习在特征空间随时间变化的场景下具有灵活优势,但仍面临三大挑战:真实数据流中混合特征类型导致传统参数建模困难;数据分布漂移引发性能骤降;时间和成本限制使全量标注不可行。为此,本文提出新算法 OL-MDISF,旨在放宽对特征类型、数据分布和监督信息的限制。方法包括:利用耦合模型构建综合潜在空间,采用自适应滑动窗口检测漂移点以保证模型稳定性,并基于几何结构关系建立标签近邻信息。通过理论分析与14个真实数据集上的实验验证,涵盖全周期评估曲线、消融研究、敏感性分析与时间集成动态,证明该方法在两类漂移场景下的有效性。本扩展版本为原始方法提供上下文分析与完整实验基准,涵盖混合特征建模、概念漂移适应与弱监督近期进展,可作为非平稳、异构、弱监督数据流研究的可复现资源。
原文摘要 · Abstract (English)
Online learning, where feature spaces can change over time, offers a flexible learning paradigm that has attracted considerable attention. However, it still faces three significant challenges. First, the heterogeneity of real-world data streams with mixed feature types presents challenges for traditional parametric modeling. Second, data stream distributions can shift over time, causing an abrupt and substantial decline in model performance. Additionally, the time and cost constraints make it infeasible to label every data instance in a supervised setting. To overcome these challenges, we propose a new algorithm Online Learning from Mix-typed, Drifted, and Incomplete Streaming Features (OL-MDISF), which aims to relax restrictions on both feature types, data distribution, and supervision information. Our approach involves utilizing copula models to create a comprehensive latent space, employing an adaptive sliding window for detecting drift points to ensure model stability, and establishing label proximity information based on geometric structural relationships. To demonstrate the model's efficiency and effectiveness, we provide theoretical analysis and comprehensive experimental results. This extension serves as a standalone technical reference to the original OL-MDISF method. It provides (i) a contextual analysis of OL-MDISF within the broader landscape of online learning, covering recent advances in mixed-type feature modeling, concept drift adaptation, and weak supervision, and (ii) a comprehensive set of experiments across 14 real-world datasets under two types of drift scenarios. These include full CER trends, ablation studies, sensitivity analyses, and temporal ensemble dynamics. We hope this document can serve as a reproducible benchmark and technical resource for researchers working on nonstationary, heterogeneous, and weakly supervised data streams.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。