研究数据预处理如何影响模型预测多样性,揭示平衡与过滤的双面作用。
Investigating the Impact of Balancing, Filtering, and Complexity on Predictive Multiplicity: A Data-Centric Perspective
- 通过21个真实数据集实验,分析平衡与过滤对预测多样性的具体影响
- 发现数据复杂度越高,预处理越可能加剧预测不一致现象
- 适合关注模型可靠性与数据质量的从业者参考
Rashomon效应在模型选择中构成重大挑战:多个模型在相同数据上表现相近但预测结果差异显著,导致预测多样性。尤其在高风险场景下,任意选择模型可能引发严重后果。传统方法仅追求准确率,忽视此问题。类别不平衡和无关变量会进一步恶化情况。数据中心人工智能通过优化数据,特别是预处理手段,可缓解该问题。然而近期研究指出,预处理可能反而加剧预测多样性。本文系统考察了平衡与过滤等预处理技术对预测多样性及模型稳定性的影响,考虑数据复杂度因素。在21个真实世界数据集上开展实验,应用多种预处理方法,并利用Rashomon效应评估其引入的预测多样性水平。同时分析过滤技术如何降低冗余并提升模型泛化能力。结果揭示了平衡方法、数据复杂度与预测多样性之间的关系,证明数据驱动策略可有效提升模型性能。
原文摘要 · Abstract (English)
The Rashomon effect presents a significant challenge in model selection. It occurs when multiple models achieve similar performance on a dataset but produce different predictions, resulting in predictive multiplicity. This is especially problematic in high-stakes environments, where arbitrary model outcomes can have serious consequences. Traditional model selection methods prioritize accuracy and fail to address this issue. Factors such as class imbalance and irrelevant variables further complicate the situation, making it harder for models to provide trustworthy predictions. Data-centric AI approaches can mitigate these problems by prioritizing data optimization, particularly through preprocessing techniques. However, recent studies suggest preprocessing methods may inadvertently inflate predictive multiplicity. This paper investigates how data preprocessing techniques like balancing and filtering methods impact predictive multiplicity and model stability, considering the complexity of the data. We conduct the experiments on 21 real-world datasets, applying various balancing and filtering techniques, and assess the level of predictive multiplicity introduced by these methods by leveraging the Rashomon effect. Additionally, we examine how filtering techniques reduce redundancy and enhance model generalization. The findings provide insights into the relationship between balancing methods, data complexity, and predictive multiplicity, demonstrating how data-centric AI strategies can improve model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。