arXiv:2501.00418cs.LGcs.AI2025-01ACL被引 7

研究大模型能否从弱模型继承可信性,发现部分属性可迁移,隐私不可。

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

  • 用正则化训练弱模型,再微调强模型以传递可信性
  • 公平性、抗攻击和分布外鲁棒性在正则化后显著提升
  • 首次探索可信性迁移,适合关注AI安全与可靠性研究者

生成式AI尤其是大语言模型的快速普及,使其广泛应用于各类场景。近年来,弱到强泛化现象(即强模型在弱模型输出上微调后性能超越弱模型)备受关注。然而,关键可信性属性如鲁棒性、公平性和隐私性是否也能实现类似迁移仍未知。本文首次系统研究弱到强可信性泛化问题:当强模型在弱模型输出上微调时,能否继承其可信性。为此提出两种基础训练策略:1)弱可信性微调(Weak TFT),在弱模型微调中引入可信性正则化;2)弱与弱到强可信性微调(Weak+WTS TFT),将正则化扩展至弱和强模型。在真实数据集上的实验表明,当双模型均正则化时,公平性、对抗鲁棒性和分布外鲁棒性均有显著提升,但隐私性未表现出弱到强可信性迁移迹象。本工作为可信性泛化提供了重要洞见。

原文摘要 · Abstract (English)

The rapid proliferation of generative AI, especially large language models, has led to their integration into a variety of applications. A key phenomenon known as weak-to-strong generalization - where a strong model trained on a weak model's outputs surpasses the weak model in task performance - has gained significant attention. Yet, whether critical trustworthiness properties such as robustness, fairness, and privacy can generalize similarly remains an open question. In this work, we study this question by examining if a stronger model can inherit trustworthiness properties when fine-tuned on a weaker model's outputs, a process we term weak-to-strong trustworthiness generalization. To address this, we introduce two foundational training strategies: 1) Weak Trustworthiness Finetuning (Weak TFT), which leverages trustworthiness regularization during the fine-tuning of the weak model, and 2) Weak and Weak-to-Strong Trustworthiness Finetuning (Weak+WTS TFT), which extends regularization to both weak and strong models. Our experimental evaluation on real-world datasets reveals that while some trustworthiness properties, such as fairness, adversarial, and OOD robustness, show significant improvement in transfer when both models were regularized, others like privacy do not exhibit signs of weak-to-strong trustworthiness. As the first study to explore trustworthiness generalization via weak-to-strong generalization, our work provides valuable insights into the potential and limitations of weak-to-strong generalization.

可信性大模型泛化微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。