arXiv:2601.05882cs.CLcs.AI2026-01中稿 · EMNLP被引 1

研究偏好调优在领域迁移下的泛化与多样性,发现伪标签可缓解性能下降但导致模式崩溃。

An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift

  • 对比五种对齐目标与多种适应策略,系统评估领域迁移表现
  • 伪标签策略显著减少领域偏移带来的性能下降,但引发输出模式单一
  • 揭示了模型泛化能力与输出多样性之间的权衡关系,适合关注对齐鲁棒性的研究者

偏好调优通过优化显式偏好信号而非仅似然性,使基础语言模型与人类对质量、帮助性或安全性的判断对齐。已有研究表明,偏好调优在训练域外性能下降且帮助性减弱。然而,现有适配策略能否缓解这种领域偏移仍不明确。本文开展系统性研究,评估不同对齐目标及从源域到目标域的多种适配策略(包括目标域监督微调与伪标签法)在摘要生成、问答帮助性及安全对齐任务中的表现。结果揭示了不同对齐目标在领域偏移下存在系统性差异。伪标签策略能显著降低领域偏移造成的性能下降,但引发模式崩溃,暴露泛化性与多样性之间的权衡。该研究为提升对齐模型的跨域鲁棒性提供了实证依据。

原文摘要 · Abstract (English)

Preference tuning aligns base language models to human judgments of quality, helpfulness, or safety by optimizing over explicit preference signals rather than likelihood alone. Prior work has shown that preference tuning degrades performance and reduces helpfulness outside the training domain. However, the extent to which adaptation strategies mitigate this domain shift remains unexplored. We address this challenge by conducting a comprehensive and systematic study of alignment generalization under domain shift. We compare five popular alignment objectives and various adaptation strategies from source to target, including target-domain supervised fine-tuning and pseudo-labeling, across summarization, question-answering helpfulness, and safety alignment tasks. Our findings reveal systematic differences in generalization across alignment objectives under domain shift. We show that adaptation strategies based on pseudo-labeling substantially reduce domain-shift degradation but induce mode collapse, revealing a generalization-diversity trade-off.

偏好调优领域迁移泛化性多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。