防止模型因恶意诱导而错误拒绝服务,确保拒绝是真不确定而非故意歧视。
Confidential Guardian: Cryptographically Prohibiting the Abuse of Model Abstention
- 通过伪造低置信度,攻击者可隐蔽歧视特定用户。
- 新攻击叫Mirage,降低目标区域置信度但不影响整体性能。
- 用零知识证明和校准检测,验证拒绝是否真实出于不确定性。
谨慎预测——即模型在不确定时选择不输出——对安全关键应用中减少有害错误至关重要。本文发现一种新型威胁:不诚实机构可利用该机制,以不确定性为名实施歧视或不合理拒服。我们提出名为Mirage的不确定性诱导攻击,故意降低特定输入区域的置信度,从而隐蔽地针对特定个体。该攻击在保持所有数据点高预测性能的同时实现歧视。为此,我们提出Confidential Guardian框架:通过分析参考数据集上的校准指标,检测人为压低的置信度;并采用零知识证明验证推理结果,确保报告的置信度真实来自部署模型。该方案防止提供方伪造置信值,同时保护模型知识产权。实验表明,Confidential Guardian能有效防范谨慎预测的滥用,提供可验证保证,确保拒绝反映真实不确定性而非恶意意图。
原文摘要 · Abstract (English)
Cautious predictions -- where a machine learning model abstains when uncertain -- are crucial for limiting harmful errors in safety-critical applications. In this work, we identify a novel threat: a dishonest institution can exploit these mechanisms to discriminate or unjustly deny services under the guise of uncertainty. We demonstrate the practicality of this threat by introducing an uncertainty-inducing attack called Mirage, which deliberately reduces confidence in targeted input regions, thereby covertly disadvantaging specific individuals. At the same time, Mirage maintains high predictive performance across all data points. To counter this threat, we propose Confidential Guardian, a framework that analyzes calibration metrics on a reference dataset to detect artificially suppressed confidence. Additionally, it employs zero-knowledge proofs of verified inference to ensure that reported confidence scores genuinely originate from the deployed model. This prevents the provider from fabricating arbitrary model confidence values while protecting the model's proprietary details. Our results confirm that Confidential Guardian effectively prevents the misuse of cautious predictions, providing verifiable assurances that abstention reflects genuine model uncertainty rather than malicious intent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。