arXiv:2604.05872cs.CRcs.AI2026-04

评测大模型在瑞士金融监管下的可靠性与安全防御能力

Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts

  • 新增两个维度,评估模型自评可靠性和对抗攻击防御力
  • 模型自评准确率73%-94%,但真实安全得分仅20%-61%
  • 适用于关注瑞士合规的AI开发者与监管机构

大型语言模型在瑞士金融与监管场景中的部署,需实证支持其生产可靠性与对抗安全性,而现有瑞士本地评估框架未同时涵盖这两方面。本文提出Swiss-Bench 003(SBP-003),将HAAS评估体系从六维扩展至八维,新增D7(自评可靠性代理)和D8(对抗安全性)。在四种语言(德、法、意、英)的808个瑞士特化题目上,对十款前沿模型进行评估,覆盖七项瑞士适配基准:Swiss TruthfulQA、Swiss IFEval、Swiss SimpleQA、Swiss NIAH、Swiss PII-Scope、System Prompt Leakage 和 Swiss German Comprehension,针对FINMA Guidance 08/2024、修订版联邦数据保护法(nDSG)及OWASP LLM Top 10风险。自评D7得分(73%-94%)显著高于外部评测的D8安全得分(20%-61%),二者评分标准不可比。系统提示泄露防御力为24.8%-88.2%,而个人身份信息(PII)提取防护普遍薄弱(14%-42%)。Qwen 3.5 Plus在自评中表现最佳(94.4%),而成本最低的GPT-oss 120B在对抗安全上最高(60.7%)。所有评估均为零样本,基于默认设置;D7为自评,非独立验证准确性。本文提供维度与FINMA、nDSG、OWASP风险分类的概念映射表。

原文摘要 · Abstract (English)

The deployment of large language models (LLMs) in Swiss financial and regulatory contexts demands empirical evidence of both production reliability and adversarial security, dimensions not jointly operationalized in existing Swiss-focused evaluation frameworks. This paper introduces Swiss-Bench 003 (SBP-003), extending the HAAS (Helvetic AI Assessment Score) from six to eight dimensions by adding D7 (Self-Graded Reliability Proxy) and D8 (Adversarial Security). I evaluate ten frontier models across 808 Swiss-specific items in four languages (German, French, Italian, English), comprising seven Swiss-adapted benchmarks (Swiss TruthfulQA, Swiss IFEval, Swiss SimpleQA, Swiss NIAH, Swiss PII-Scope, System Prompt Leakage, and Swiss German Comprehension) targeting FINMA Guidance 08/2024, the revised Federal Act on Data Protection (nDSG), and OWASP Top 10 for LLMs. Self-graded D7 scores (73-94%) exceed externally judged D8 security scores (20-61%) by a wide margin, though these dimensions use non-comparable scoring regimes. System prompt leakage resistance ranges from 24.8% to 88.2%, while PII extraction defense remains weak (14-42%) across all models. Qwen 3.5 Plus achieves the highest self-graded D7 score (94.4%), while GPT-oss 120B achieves the highest D8 score (60.7%) despite being the lowest-cost model evaluated. All evaluations are zero-shot under provider default settings; D7 is self-graded and does not constitute independently validated accuracy. I provide conceptual mapping tables relating benchmark dimensions to FINMA model validation requirements, nDSG data protection obligations, and OWASP LLM risk categories.

大模型评测瑞士合规对抗安全数据隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。