测试发现大模型对同义句式敏感,评分波动大,评估需加强鲁棒性。
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
- 用同义词替换和句法变换生成语义不变的输入
- 近半数模型在同义句下性能下降超10%,复杂任务更明显
- 模型越大越稳定不成立,不同任务差异显著
大语言模型(LLMs)的快速进步使标准化评估基准成为模型比较的主要工具。然而,由于对输入提示的微小变化敏感,其可靠性日益受到质疑。本文研究了受控的、语义等价的词汇和句法扰动对23个主流LLMs在MMLU、SQuAD和AMEGA三个基准上的绝对表现和相对排名的影响。我们采用两种语言学严谨的流水线生成语义保持的变体:一是进行同义词替换实现词汇变化,二是利用依存句法分析确定适用的句法转换。结果表明,词汇扰动几乎在所有模型和任务中均导致显著的性能下降,统计上显著;而句法扰动影响更异质,偶尔提升成绩。两类扰动均使复杂任务上的模型排行榜不稳定。此外,模型鲁棒性并未随模型规模一致提升,表现出强任务依赖性。总体而言,结果表明LLMs更依赖表层词汇模式而非抽象语言能力,强调将鲁棒性测试作为大模型评估的标配。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) has established standardized evaluation benchmarks as the primary instrument for model comparison. Yet, their reliability is increasingly questioned due to sensitivity to shallow variations in input prompts. This paper examines how controlled, truth-conditionally equivalent lexical and syntactic perturbations affect the absolute performance and relative ranking of 23 contemporary LLMs across three benchmarks: MMLU, SQuAD, and AMEGA. We employ two linguistically principled pipelines to generate meaning-preserving variations: one performing synonym substitution for lexical changes, and another using dependency parsing to determine applicable syntactic transformations. Results show that lexical perturbations consistently induce substantial, statistically significant performance degradation across nearly all models and tasks, while syntactic perturbations have more heterogeneous effects, occasionally improving results. Both perturbation types destabilize model leaderboards on complex tasks. Furthermore, model robustness did not consistently scale with model size, revealing strong task dependence. Overall, the findings suggest that LLMs rely more on surface-level lexical patterns than on abstract linguistic competence, underscoring the need for robustness testing as a standard component of LLM evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。