arXiv:2605.29816cs.AI2026-05

通过微调消除模型偏差,提升大模型对提示词变化的鲁棒性。

Harnessing non-adversarial robustness in large language models

论文配图:Harnessing non-adversarial robustness in large language models
图 1 · 摘自论文原文
  • 提出去偏微调方法,无需重训整个模型即可增强鲁棒性。
  • 实验证明该方法能有效应对随机提示扰动,提升模型稳定性。
  • 适合关注模型可靠性与安全性研究者使用。

本文针对大语言模型在语义相似但文本不同的提示词下表现不稳定的问题,提出一种无需昂贵全模型重训的鲁棒性增强方法。理论分析揭示神经网络模块输出存在系统性偏移或扰动诱导偏差,是影响鲁棒性的关键因素。基于此,提出‘为鲁棒性而进行的去偏’微调策略。研究识别了去偏有效的条件,并通过理论与大量实验验证,该方法可快速高效地提升模型对随机提示扰动的鲁棒性,具备提供鲁棒性认证的潜力。

原文摘要 · Abstract (English)

The work presents an approach for addressing the challenge of robustness in Large Language Models (LLMs) to alterations and potential errors caused by semantically similar but textually different prompts. Recent works have shown that these kinds of prompt variations can significantly impact the performance of LLMs on tasks. The central question is: can LLMs' robustness to semantically-neutral prompt alterations be acquired without expensive retraining of the entire model? We address this question both theoretically and through experiments. Our theoretical analysis reveals a crucial factor impacting model robustness - a systematic expected shift or perturbation-induced bias in neural network module outputs. Motivated by this analysis, we show that robustness can be achieved via a simple fine-tuning process: debiasing for robustness. We identify conditions when debiasing helps and when it does not, and demonstrate, through both theory and extensive experiments, that debiasing for robustness may indeed be a quick and efficient tool to enhance robustness and provide certification against random prompt perturbations.

大模型鲁棒性微调去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。