arXiv:2607.20558cs.LGcs.AI2026-07

测试大模型在多轮对话中的稳定性,发现性能普遍下降。

StabilityBench: Benchmarking Instability in LLMs

论文配图:StabilityBench: Benchmarking Instability in LLMs
图 1 · 摘自论文原文
  • 将单轮测试转为多轮交互,模拟真实用户行为
  • 九个模型在四类任务中均出现显著性能下降
  • 适合关注模型鲁棒性的研究人员和开发者

AI助手在医疗、政府等高风险场景中日益普及,但其实际表现因上下文依赖性强而难以理解。现有评估多采用静态、单轮方式,难以捕捉真实对话的变异性。本文提出StabilityBench,一种通用、模型无关的基准增强框架,将单轮测试转化为多轮交互历史,通过人口统计代理或奉承性诱导注入真实用户行为,同时保持原任务意图。我们在数学推理、健康问答与安全四个基准上,评估了九个大语言模型。结果表明,模型在注入后性能普遍不稳定,三类任务出现明显下降。研究揭示静态评估的局限性,推动更真实的测试环境。为此,我们提出StabilityBench-Mini,在不增加成本的前提下,通过跨多样化维度采样实现更贴近现实的评估。

原文摘要 · Abstract (English)

AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poorly understood due to strong context dependence. Current evaluation protocols follow a defense-in-depth paradigm with compounding layers of safeguards, ranging from traditional benchmarks to live or adversarial testing. Such benchmarks remain largely static and single-turn, limiting their ability to capture real-world variability in conversational settings. We propose StabilityBench, a principled, general and model-agnostic benchmark operator that transforms single-turn benchmark queries into multi-turn interaction histories. StabilityBench augments existing benchmarks by injecting realistic user simulations, through demographic proxies or sycophantic baits, while preserving original task intent. We apply StabilityBench to four benchmarks spanning mathematical reasoning, health question-answering and safety, and evaluate nine large language models under these conditions. Our results show that model performance is consistently unstable under these injections, with considerable performance degradations on three out of four benchmarks studied. These highlight important limitations of static evaluations and motivate more realistic evaluation settings. To this end, we propose StabilityBench-Mini: a size-preserving variant of StabilityBench that samples across diversification axes, enabling more realistic evaluation without increasing costs.

大模型评估稳定性测试多轮对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。