用模糊逻辑提升大数据模型可解释性,兼顾公平与效率。
A Fuzzy Logic-Based Framework for Explainable Machine Learning in Big Data Analytics
- 结合二型模糊集与聚类,增强模型对噪声数据的解释能力。
- 在空气质量数据上实现4%的轮廓系数提升,熵值降低1%。
- 适合需要透明决策的环保、医疗等高风险领域应用。
大数据分析中机器学习模型日益复杂,尤其在环境监测等领域,可解释性对建立信任、符合伦理和法规(如GDPR)至关重要。传统黑箱模型缺乏透明度,而事后解释方法(如LIME、SHAP)常牺牲准确率或无法提供内在洞察。本文提出一种融合二型模糊集、粒计算与聚类的新框架,以提升大数据环境下的可解释性与公平性。在UCI空气质量数据集上的实验表明,该框架有效处理传感器数据的不确定性,生成语言规则,并通过轮廓系数与熵值评估公平性。主要贡献包括:(1) 二型模糊聚类使凝聚度较一型方法提升约4%(轮廓系数0.365 vs. 0.349),公平性熵值达0.918;(2) 引入公平性度量,缓解无监督场景中的偏差;(3) 基于规则的组件实现内在XAI,平均覆盖率0.65;(4) 可扩展评估显示线性运行时间(采样大数据规模下约0.005秒)。实验结果表明,相比DBSCAN与层次聚类等基线方法,该方法在可解释性、公平性和效率方面均表现更优,轮廓系数提升4%,公平性熵减少最高达1%。
原文摘要 · Abstract (English)
The growing complexity of machine learning (ML) models in big data analytics, especially in domains such as environmental monitoring, highlights the critical need for interpretability and explainability to promote trust, ethical considerations, and regulatory adherence (e.g., GDPR). Traditional "black-box" models obstruct transparency, whereas post-hoc explainable AI (XAI) techniques like LIME and SHAP frequently compromise accuracy or fail to deliver inherent insights. This paper presents a novel framework that combines type-2 fuzzy sets, granular computing, and clustering to boost explainability and fairness in big data environments. When applied to the UCI Air Quality dataset, the framework effectively manages uncertainty in noisy sensor data, produces linguistic rules, and assesses fairness using silhouette scores and entropy. Key contributions encompass: (1) A type-2 fuzzy clustering approach that enhances cohesion by about 4% compared to type-1 methods (silhouette 0.365 vs. 0.349) and improves fairness (entropy 0.918); (2) Incorporation of fairness measures to mitigate biases in unsupervised scenarios; (3) A rule-based component for intrinsic XAI, achieving an average coverage of 0.65; (4) Scalable assessments showing linear runtime (roughly 0.005 seconds for sampled big data sizes). Experimental outcomes reveal superior performance relative to baselines such as DBSCAN and Agglomerative Clustering in terms of interpretability, fairness, and efficiency. Notably, the proposed method achieves a 4% improvement in silhouette score over type-1 fuzzy clustering and outperforms baselines in fairness (entropy reduction by up to 1%) and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。