arXiv:2510.06125cs.LG2025-10被引 2

压缩模型未必忠实原模型,新方法能发现隐藏的偏差。

Downsized and Compromised?: Assessing the Faithfulness of Model Compression

  • 用模型一致性和卡方检验评估压缩后行为变化
  • 高准确率下仍存在显著预测模式偏移
  • 适合医疗、金融等对公平性要求高的场景

实际应用中,计算资源限制常需通过模型压缩将大模型转为小型高效版本。尽管压缩旨在降低规模和计算成本而不损失性能,但传统评估仅关注大小与准确率的权衡,忽略了模型忠实度问题。在医疗、金融、司法等高风险领域,压缩模型必须保持与原模型行为一致。本文提出一种新方法,引入一组忠实度度量,捕捉压缩前后模型行为的变化。我们使用模型一致性分析预测一致性,并应用卡方检验检测整体数据集及人口子群体中预测模式的统计显著变化,揭示了聚合公平性指标可能掩盖的偏移。通过在三个具有社会意义的数据集上对人工神经网络进行量化和剪枝,我们发现:高准确率不等于高忠实度,标准指标如准确率和等机会无法察觉细微但重要的变化。所提度量提供了一种更直接、实用的手段,确保压缩带来的效率提升不损害可信AI所需的公平性与忠实度。

原文摘要 · Abstract (English)

In real-world applications, computational constraints often require transforming large models into smaller, more efficient versions through model compression. While these techniques aim to reduce size and computational cost without sacrificing performance, their evaluations have traditionally focused on the trade-off between size and accuracy, overlooking the aspect of model faithfulness. This limited view is insufficient for high-stakes domains like healthcare, finance, and criminal justice, where compressed models must remain faithful to the behavior of their original counterparts. This paper presents a novel approach to evaluating faithfulness in compressed models, moving beyond standard metrics. We introduce and demonstrate a set of faithfulness metrics that capture how model behavior changes post-compression. Our contributions include introducing techniques to assess predictive consistency between the original and compressed models using model agreement, and applying chi-squared tests to detect statistically significant changes in predictive patterns across both the overall dataset and demographic subgroups, thereby exposing shifts that aggregate fairness metrics may obscure. We demonstrate our approaches by applying quantization and pruning to artificial neural networks (ANNs) trained on three diverse and socially meaningful datasets. Our findings show that high accuracy does not guarantee faithfulness, and our statistical tests detect subtle yet significant shifts that are missed by standard metrics, such as Accuracy and Equalized Odds. The proposed metrics provide a practical and more direct method for ensuring that efficiency gains through compression do not compromise the fairness or faithfulness essential for trustworthy AI.

模型压缩忠实度评估公平性检测统计检验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。