提出多层级扰动框架,揭示大模型在不同故障类型下的真实脆弱性。
Measuring Stability and Failure Behavior in Language Models Under Structured Perturbations

- 设计七类结构化扰动,按严重等级测试模型稳定性
- 发现冲突指令和不可能前提导致普遍崩溃,而拒答能力参差不齐
- 适合评估模型鲁棒性,尤其关注拒绝回答的可靠性
语言模型通常仅以单一准确率评估,无法反映输入扰动下的性能退化。本文提出一种分级、多家族、故障感知的应力测试框架,对同一组100个种子问题进行扩展,生成4,473个有效性验证后的测试用例。测试覆盖七类扰动:答案不变、改写、输入噪声、格式变化、无关上下文、上下文负载、矛盾指令,以及知识边界(移除可答性,使拒绝成为正确响应)。每项测试标注严重度,模型表现通过各层级准确率、加权稳定性及各家族崩溃点(相对于自身基线)总结。在四个能力层级的模型上运行后发现,失败临界点因扰动类型而异,而非全局统一;冲突指令与基于不可能前提的问题暴露了所有模型的共性弱点;对不可答性的识别在缺失信息和虚构证据下可靠,但在不可能前提下表现薄弱。这些故障点在传统准确率报告中完全不可见。
原文摘要 · Abstract (English)
Language models are usually judged by a single accuracy score, which does not reveal how their performance degrades as inputs are perturbed. We present a graded, multi-family, failure-aware framework for stress-testing reasoning models. It perturbs each problem along a multi-level severity ladder across seven families: six that preserve the answer, paraphrase, input noise, formatting, irrelevant context, context load, and conflicting instructions, and a Knowledge Boundary family that removes answerability so that refusal becomes the correct response. Every test is validity-gated and labeled by its measured severity, and each model is summarized by per-level Accuracy, a magnitude-weighted Stability, and a per-family Collapse Point defined relative to the model's own baseline. Instantiated on the same 100 seed problems used by GSM-Symbolic, expanded into 4,473 gated tests and run on four models spanning capability tiers, the framework exposes structure that an aggregate score hides: the level at which a model fails is family-specific rather than global, and two stressors expose consistent weaknesses across all models: conflicting instructions and questions built on an impossible premise. Recognition of unanswerability is otherwise uneven, reliable on missing information and fabricated evidence but weak on impossible premises. These failure points are invisible to standard accuracy reporting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。