测试大模型在企业应用中对输入微小变化的稳定性,发现性能可下降40个百分点
Evaluating Robustness of Large Language Models in Enterprise Applications: Benchmarks for Perturbation Consistency Across Formats and Languages
- 构建多类型扰动基准,涵盖格式、语言、位置等变化
- 11个模型在关键指标上性能最高下降40个百分点
- 8B模型表现反超多数更大模型,挑战规模即能力假设
企业级大模型应用需在多样场景中保持高质量与可靠性能,对输入微小变化具备鲁棒性至关重要。现有研究虽表明微小提示变化会导致输出显著差异,但主要集中在有限扰动类型和小规模学术数据集,难以反映真实应用场景。为此,我们提出一个全面的基准套件,评估多种扰动类型下的鲁棒性,包括通用文本编辑(如标点、空格)、格式变更(如JSON、YAML)、多语言与跨语言输入,以及指令位置变化。评估了11个参数量从4B到120B+的大模型,发现微小扰动导致关键企业指标性能最高下降40个百分点。关键发现是:模型规模与鲁棒性的关系比传统认知更复杂——一个8B模型(Ministral 3 8B)表现优于多数更大模型,而另一个8B模型(Llama 3.1 8B)则整体最差。
原文摘要 · Abstract (English)
Enterprise LLM applications require consistently high quality and reliable performance across diverse scenarios, demanding robustness to minor variations. Existing research shows that even small prompt changes can lead to substantial differences in output, but has mainly focused on a narrow set of perturbations with small academic datasets, limiting their relevance to real-world applications. To address this, we present a comprehensive benchmark suite that evaluates robustness across multiple perturbation types, including general text edits (e.g., punctuation, whitespace), formatting changes (e.g., JSON, YAML), multilingual and cross-lingual inputs, and positional variations in instructions. Evaluating 11 models ranging from 4B to 120B+ parameters, we find that minor perturbations reduce performance by up to 40 percentage points on key enterprise metrics. Critically, we demonstrate that the relationship between model size and robustness is more nuanced than conventional assumptions suggest: an 8B parameter model (Ministral 3 8B) outperforms most larger models, while another 8B model (Llama 3.1 8B) performs worst overall.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。