arXiv:2504.17723cs.LG2025-04被引 2

用统计方法快速验证大模型运行时鲁棒性,效率提升百倍

Statistical Runtime Verification for LLMs via Robustness Estimation

  • 通过分析置信度分布评估语义扰动下的模型鲁棒性
  • 相比传统方法验证时间从小时级降至分钟级,误差低于1%
  • 适合在无源码环境下实时监控大模型安全运行

对抗鲁棒性验证对保障大语言模型在关键应用中的安全部署至关重要。然而,由于计算开销呈指数级增长且需白盒访问,现有形式化验证方法对现代大模型不适用。本文将RoMA统计验证框架进行适配与扩展,探索其作为黑盒部署场景下在线运行时鲁棒性监测器的可行性。该方法通过分析语义扰动下的置信度分布,提供带统计保证的量化鲁棒性评估。实验对比表明,RoMA在准确率上与形式化验证基线相差不超过1%,验证时间由数小时缩短至数分钟。我们在语义、类别和拼写扰动等多个领域进行了评估,结果证明该框架在实际大模型部署中具有有效性。这些发现表明,当形式化方法不可行时,RoMA可作为一种潜在可扩展的替代方案,为基于大模型的系统运行时验证带来新可能。

原文摘要 · Abstract (English)

Adversarial robustness verification is essential for ensuring the safe deployment of Large Language Models (LLMs) in runtime-critical applications. However, formal verification techniques remain computationally infeasible for modern LLMs due to their exponential runtime and white-box access requirements. This paper presents a case study adapting and extending the RoMA statistical verification framework to assess its feasibility as an online runtime robustness monitor for LLMs in black-box deployment settings. Our adaptation of RoMA analyzes confidence score distributions under semantic perturbations to provide quantitative robustness assessments with statistically validated bounds. Our empirical validation against formal verification baselines demonstrates that RoMA achieves comparable accuracy (within 1\% deviation), and reduces verification times from hours to minutes. We evaluate this framework across semantic, categorial, and orthographic perturbation domains. Our results demonstrate RoMA's effectiveness for robustness monitoring in operational LLM deployments. These findings point to RoMA as a potentially scalable alternative when formal methods are infeasible, with promising implications for runtime verification in LLM-based systems.

大模型安全运行时验证统计检测鲁棒性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。