arXiv:2505.23933cs.LGcs.AI2025-05

用结构匹配让小模型教会大模型做人,提升AI行为对齐效果

BIRD: Behavior Induction via Representation-structure Distillation

  • 通过匹配师生模型内部表征结构实现行为迁移
  • 在图像分类中提升鲁棒性准确率最高达16%以上
  • 小教师模型也能有效指导大模型,适合安全AI部署

具备人类价值观一致行为的深度学习模型具有鲁棒性、公平性和诚实性等特性。然而将这些行为属性迁移到不同任务或数据分布的模型上仍具挑战:对齐行为在微调过程中容易丢失,且收集保留此类行为的任务特定数据成本高昂。本文提出BIRD(基于表征-结构蒸馏的行为诱导),一种灵活的框架,通过匹配学生模型与教师模型的内部表征结构来传递对齐行为。应用于图像分类中的分布外鲁棒性时,BIRD优于微调、迁移学习和持续学习方法,鲁棒准确率相较最强基线最高提升16%。即使教师模型训练于更简单数据集且体积仅为学生模型的1/25,效果依然显著。在超过400组师生对的大规模研究中,发现教师表征的三个可解释且可计算的属性——任务相关性、行为相关性和互补知识——可解释高达85%的迁移成功率方差。这些发现为教师选择与设计提供实用指导。BIRD使小型优质对齐模型成为可扩展的对齐种子,破解了安全AI系统野外部署的关键瓶颈。

原文摘要 · Abstract (English)

Human-aligned deep learning models exhibit behaviors consistent with human values, such as robustness, fairness, and honesty. Transferring these behavioral properties to models trained on different tasks or data distributions remains challenging: aligned behavior is easily forgotten during fine-tuning, and collecting task-specific data that preserves this behavior can be prohibitively costly. We introduce BIRD (Behavior Induction via Representation-structure Distillation), a flexible framework for transferring aligned behavior by matching the internal representation structure of a student model to that of a teacher. Applied to out-of-distribution robustness in image classification, BIRD outperforms fine-tuning, transfer learning, and continual learning methods, improving robust accuracy by up to 16% over the next strongest baseline. It remains effective even when the teacher is trained on a much simpler dataset and is $25 \times$ smaller than the student. In a large-scale study of over 400 teacher-student pairs, we show that three interpretable and computable properties of the teacher's representations (i.e., task relevance, behavioral relevance, and complementary knowledge) explain up to 85% of the variance in transfer success. These insights offer practical guidance for teacher selection and design. BIRD turns small, well-aligned models into scalable alignment seeds, removing a key bottleneck in deploying safe AI systems in the wild.

行为对齐模型蒸馏鲁棒性AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。