研究发现,缓解性别刻板印象会损害模型下游性能,且现有方法难以兼顾。
Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
- 通过控制遗忘机制评估公平性目标的有效性
- 大规模数据集下整体遗忘程度与任务性能强相关
- 当前方法无法同时减少刻板印象和整体遗忘,适合关注对齐风险的研究者
道德对齐已成为调控预训练语言模型行为的常用方法,通常通过在精选数据集上微调实现。性别刻板印象缓解是其代表性应用之一。然而,该过程常伴随下游任务性能下降。以往研究多通过设计公平性目标,引导模型选择性遗忘刻板知识,同时保留语言建模能力(即整体遗忘)。本文从遗忘视角分析这一性能权衡是否可实现。结果表明:(1)下游任务性能与整体遗忘程度强相关;(2)选择性遗忘虽能降低刻板印象,但整体遗忘随之增加;(3)现有缓解遗忘的通用方案对减少整体遗忘无效,也无法提升下游性能。
原文摘要 · Abstract (English)
Moral alignment has emerged as a widely adopted approach for regulating the behavior of pretrained language models (PLMs), typically through fine-tuning on curated datasets. Gender stereotype mitigation is a representational task within the broader application of moral alignment. However, this process often comes at the cost of degraded downstream task performance. Prior studies commonly aim to achieve a performance trade-off by encouraging PLMs to selectively forget only stereotypical knowledge through carefully designed fairness objective, while preserving their language modeling capability (overall forgetting). In this short paper, we investigate whether the performance trade-off can be achieved through the lens of forgetting and the fairness objective. Our analysis shows that the large datasets needed for satisfactory fairness highlight the limitations of current fairness objectives in achieving an effective trade-off: (1) downstream task performance is strongly correlated with overall forgetting; (2) selective forgetting reduces stereotypes, but overall forgetting increases. and (3) general solutions for alleviating forgetting are ineffective at reducing the overall forgetting and fail to improve downstream task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。