低秩压缩让大模型更省资源,但隐私与公平性受影响。
Decomposed Trust: Privacy, Adversarial Robustness, Ethics, and Fairness in Low-Rank LLMs
- 通过低秩分解压缩模型,降低计算和内存开销。
- 压缩后对话中个人身份信息保护变弱,公平性下降。
- 揭示了模型各层对对抗鲁棒性的贡献机制。
大语言模型(LLMs)推动了多个领域的进步,但其庞大体量限制了在资源受限环境中的部署。低秩分解通过压缩模型有效减少计算与内存消耗,同时保持精度。尽管这些压缩模型表现出良好的性能与系统优势,其可信度影响仍不明确。本文首次全面研究低秩分解对LLM在隐私、对抗鲁棒性、伦理与公平性方面的影响,并通过可解释性分析揭示内在机制。我们评估了多种大小和架构的LLM在不同低秩分解算法下的表现,发现:(1)训练数据隐私得以保留,但对话中个人身份信息保护减弱;(2)压缩普遍增强对抗鲁棒性;(3)零样本提示下伦理性能下降,少样本提示下部分恢复;(4)公平性在压缩后下降。此外,研究还探讨了模型规模与微调对可信度的影响。为突破黑箱分析,我们采用基于梯度的归因方法,识别出对对抗鲁棒性贡献最大的模型层。
原文摘要 · Abstract (English)
Large language models (LLMs) have driven major advances across domains, yet their massive size hinders deployment in resource-constrained settings. Low-rank factorization addresses this challenge by compressing models to effectively reduce their computation and memory consumption while maintaining accuracy. While these compressed models boast benign performance and system-level advantages, their trustworthiness implications remain poorly understood. In this paper, we present the first comprehensive study of how low-rank factorization affects LLM trustworthiness across privacy, adversarial robustness, ethics, and fairness, complemented by an explainability-driven analysis of the internal mechanisms behind these trust-related changes. We evaluate multiple LLMs of different sizes and architectures compressed with various low-rank factorization algorithms, revealing key insights: (1) low-rank factorization preserves training data privacy but weakens the protection of personally identifiable information during conversations; (2) adversarial robustness is generally enhanced under compression; (3) ethics degrades in zero-shot prompting but partially recovers in few-shot prompting; (4) fairness declines under compression. Beyond compression, we investigate how model scale and fine-tuning affect trustworthiness. Additionally, to move beyond black-box analysis, we employ a gradient-based attribution to identify which layers of LLMs contribute most to adversarial robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。