测试大模型在噪声和变形任务下的表现,发现不同规模模型稳定性差异。
LANTERN: Language Model Assessment on Noisy and Transformed Tasks for Understanding Error and Robustness Nuances

- 构建合成数据集,模拟多种输入扰动评估鲁棒性
- 小中大模型在扰动下响应变异显著,指令微调模型更敏感
- 为提升模型可靠性与评测方法提供实证依据
大型语言模型(LLMs)的鲁棒性评估仍是关键挑战,尤其在输入数据存在扰动时。本文系统评估了多维度下的模型鲁棒性,包括词错误率、字符重复与重复、选项修改及指令遵循变异性。为此,我们构建了一个涵盖多种选择题(MCQ)数据集和指令跟随任务的合成增强数据集。在不同规模(小、中、大)以及基础与指令微调版本的LLM上进行了广泛实验。分析量化了模型在扰动条件下的响应变化,并揭示了与基线模型的差异。结果为理解不同评估场景下模型的稳定性提供了洞见,有助于开发更鲁棒可靠的LLM及评估方法。
原文摘要 · Abstract (English)
Robustness evaluation of large language models (LLMs) remains a critical challenge, particularly in assessing their sensitivity to perturbations in input data. In this work, we systematically evaluate LLM robustness across multiple dimensions, including word error rate, character repetition and duplication, modifications in choices, and variability in instruction following. To facilitate this evaluation, we construct a synthetic and augmented dataset encompassing a diverse set of LLM benchmarks, specifically targeting multiple-choice question (MCQ) datasets and instruction-following tasks. We conduct extensive experiments on LLMs of varying scales-small, medium, and large-as well as across base and instruction-tuned variants. Our analysis quantifies the variability in model responses under perturbed conditions and highlights discrepancies relative to baseline models. The findings provide insights into the stability of LLMs across different evaluation scenarios contributing to the development of more robust and reliable language models as well as robust evaluation methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。