arXiv:2512.17460cs.SEcs.LG2025-12被引 2

首次系统研究缺陷预测中数据问题的共现影响,揭示真实数据的复杂性。

When Data Quality Issues Collide: A Large-Scale Empirical Study of Co-Occurring Data Quality Issues in Software Defect Prediction

  • 在374个数据集上同时分析五类数据问题的交互作用
  • 发现93%以上数据集存在多种问题共现,关键阈值明确
  • 提出需结合上下文评估模型表现,避免单一方法失效

软件缺陷预测(SDP)模型是主动保障软件质量的核心,但其效果常受限于数据质量。以往研究多孤立分析类别不平衡或特征无关等问题,忽视了真实数据中多个问题常同时出现且相互影响。本研究首次在374个数据集和五种分类器上,大规模实证分析五类共现数据质量问题(类别不平衡、类别重叠、无关特征、属性噪声、异常值)的协同效应。采用可解释提升机与分层交互分析,量化默认参数下直接与条件影响,反映实际使用基准。结果表明,共现几乎普遍存在:即使最罕见的问题(属性噪声)也出现在超过93%的数据集中。无关特征与不平衡近乎普遍,而类别重叠是最具破坏性的因素。我们识别出稳定临界点:类别重叠约0.20,不平衡0.65–0.70,无关特征0.94,超过即多数模型性能下降。还发现反直觉现象,如当无关特征较小时,异常值反而提升性能,强调需上下文敏感评估。最终揭示性能与鲁棒性间的权衡:无单一学习器在所有条件下占优。通过联合分析问题频次、共现模式、阈值及条件效应,本研究填补了SDP研究中长期存在的空白,推动从孤立分析转向全面、数据驱动的理解。

原文摘要 · Abstract (English)

Software Defect Prediction (SDP) models are central to proactive software quality assurance, yet their effectiveness is often constrained by the quality of available datasets. Prior research has typically examined single issues such as class imbalance or feature irrelevance in isolation, overlooking that real-world data problems frequently co-occur and interact. This study presents, to our knowledge, the first large-scale empirical analysis in SDP that simultaneously examines five co-occurring data quality issues (class imbalance, class overlap, irrelevant features, attribute noise, and outliers) across 374 datasets and five classifiers. We employ Explainable Boosting Machines together with stratified interaction analysis to quantify both direct and conditional effects under default hyperparameter settings, reflecting practical baseline usage. Our results show that co-occurrence is nearly universal: even the least frequent issue (attribute noise) appears alongside others in more than 93% of datasets. Irrelevant features and imbalance are nearly ubiquitous, while class overlap is the most consistently harmful issue. We identify stable tipping points around 0.20 for class overlap, 0.65-0.70 for imbalance, and 0.94 for irrelevance, beyond which most models begin to degrade. We also uncover counterintuitive patterns, such as outliers improving performance when irrelevant features are low, underscoring the importance of context-aware evaluation. Finally, we expose a performance-robustness trade-off: no single learner dominates under all conditions. By jointly analyzing prevalence, co-occurrence, thresholds, and conditional effects, our study directly addresses a persistent gap in SDP research. Hence, moving beyond isolated analyses to provide a holistic, data-aware understanding of how quality issues shape model performance in real-world settings.

缺陷预测数据质量共现问题实证研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。