arXiv:2505.05704cs.CLcs.AI2025-05被引 6

对比三种微调方法在虚假相关下的表现,发现无一通用最优解。

Assessing Robustness to Spurious Correlations in Post-Training Language Models

  • 设计合成任务测试SFT、DPO、KTO对虚假相关性的鲁棒性
  • 90%虚假相关时模型性能普遍下降,但数学推理中偏好方法更稳
  • 复杂任务选SFT,需根据任务类型选择微调策略

监督微调和基于偏好的微调已成为对齐大语言模型与用户意图及正确性标准的主流方法。然而,真实训练数据常存在虚假相关性——源于偏差、数据集特征或其它“捷径”特征——可能损害模型性能或泛化能力。本文系统评估了三种后训练算法:监督微调(SFT)、直接偏好优化(DPO)和KTO(Kahneman-Tversky优化),覆盖数学推理、受限指令遵循和文档问答等合成任务,以及不同虚假相关程度(10% vs. 90%)和两类数据特征(特征模糊性与分布狭窄性)。结果表明,模型在高虚假相关下通常但并非总是退化;偏好方法在数学推理中表现出相对鲁棒性,而SFT在复杂上下文任务中保持更强性能。研究揭示:无单一方法在所有场景中占优,最佳选择取决于目标任务类型与虚假相关性质。

原文摘要 · Abstract (English)

Supervised and preference-based fine-tuning techniques have become popular for aligning large language models (LLMs) with user intent and correctness criteria. However, real-world training data often exhibits spurious correlations -- arising from biases, dataset artifacts, or other "shortcut" features -- that can compromise a model's performance or generalization. In this paper, we systematically evaluate three post-training algorithms -- Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and KTO (Kahneman-Tversky Optimization) -- across a diverse set of synthetic tasks and spuriousness conditions. Our tasks span mathematical reasoning, constrained instruction-following, and document-grounded question answering. We vary the degree of spurious correlation (10% vs. 90%) and investigate two forms of artifacts: "Feature Ambiguity" and "Distributional Narrowness." Our results show that the models often but not always degrade under higher spuriousness. The preference-based methods (DPO/KTO) can demonstrate relative robustness in mathematical reasoning tasks. By contrast, SFT maintains stronger performance in complex, context-intensive tasks. These findings highlight that no single post-training strategy universally outperforms in all scenarios; the best choice depends on the type of target task and the nature of spurious correlations.

大模型微调虚假相关鲁棒性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。