任务向量可调优以平衡准确率与公平性,适合高风险场景的模型编辑。
On Fairness of Task Arithmetic: The Role of Task Vectors
- 通过调节任务向量实现高效模型编辑,避免全量微调。
- 在多个模型和数据集上,公平性指标提升12%-23%。
- 支持按子群体定制向量,适配需要公平性的应用。
模型编辑技术,特别是基于任务向量的任务算术,通过简单的参数运算实现直接更新,为全量微调提供高效的替代方案。尽管该方法显著降低计算开销,其对公平性的影响仍缺乏系统研究,尤其在仇恨言论检测等高风险应用场景中。本文首次在二分类文本与图像任务中系统评估了任务算术的群体公平性,对比全量微调(FFT)和低秩适应(LoRA)。我们在多个语言模型与数据集上采用标准公平性度量(如人口均等性与机会均等性)进行评估。结果表明,通过调整任务向量可在保持竞争力准确率的同时降低偏差;将子群体特定任务向量合并可有效引导公平性结果。此外,我们提供了任务向量缩放与公平性指标间的理论界,揭示了观察到的权衡关系。这些发现确立了任务算术不仅是一种低成本编辑方法,更是在标准群体公平分类设置下具备公平意识的替代方案,为大模型负责任部署奠定基础。
原文摘要 · Abstract (English)
Model editing techniques, particularly task arithmetic with task vectors, offer an efficient alternative to full fine-tuning by enabling direct parameter updates through simple arithmetic operations. While this approach promises substantial computational savings, its impact on fairness has remained largely unexplored -- despite growing concern over biased outcomes in high-stakes applications such as hate speech detection. In this work, we present the first systematic study of group fairness in task arithmetic within this binary text and image classification regime, comparing it against full fine-tuning (FFT) and Low-Rank Adaptation (LoRA). We evaluate across multiple language models and datasets using standard group fairness metrics, including Demographic Parity and Equalized Odds. Our analysis shows that task vectors can be tuned to achieve competitive accuracy while reducing disparities, and that merging subgroup-specific task vectors provides a practical mechanism for steering fairness outcomes. We further provide a theoretical bound linking task vector scaling to fairness metrics, offering insight into the observed trade-offs. Together, these findings establish task arithmetic not only as a cost-efficient editing method but also as a fairness-aware alternative to existing adaptation techniques, within the standard group-fair classification setting, laying the groundwork for responsible deployment of large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。