arXiv:2603.10807q-fin.CPcs.AI2026-03

为金融领域大模型安全评估设计风险敏感评分,提升真实场景下的漏洞检测能力。

Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services

  • 提出风险调整伤害评分RAHS,综合严重性、免责声明和评委一致性
  • 在989个金融场景提示中,多轮攻击暴露更严重的合规漏洞
  • 适合金融AI安全测试、监管合规团队使用

现有大模型安全评估依赖二元攻击成功率和通用分类体系,导致金融、保险等受监管领域部署存在法律或专业合理框架下的失效风险。本文提出RAHS(风险调整伤害评分),一种融合披露严重性、免责声明影响及评委一致性的风险敏感指标,并构建FinRedTeamBench基准,包含989个提示,覆盖7个金融风险领域与34个子类别,映射至监管框架。评估采用三个异构大模型组成的集成判断系统,经人类专家验证,并结合自适应多轮红队测试流程。在九个开源权重模型上,RAHS在接近满成功率条件下仍保持排序稳定性,多轮压力测试不仅提升越狱率,更触发更严重的实际披露行为,揭示了单轮、通用评估无法发现的失效模式。

原文摘要 · Abstract (English)

Existing LLM safety evaluations rely on binary attack-success rates and domain-agnostic taxonomies, leaving regulated Banking, Financial Services, and Insurance (BFSI) deployments exposed to failures elicited through legally or professionally plausible framing. We introduce RAHS (Risk-Adjusted Harm Score), a risk-sensitive metric jointly capturing disclosure severity, disclaimer mitigation, and inter-judge agreement, and FinRedTeamBench, a 989-prompt benchmark spanning seven BFSI risk areas and 34 sub-categories mapped to regulatory frameworks. Evaluation uses an ensemble of three heterogeneous LLM judges, validated against human experts, and an adaptive multi-turn red-teaming pipeline. On nine open-weight models, RAHS preserves separation under near-ceiling ASR, ranking is stable under hyperparameter sweeps, and multi-turn pressure drives not only more jailbreaks but more operationally severe disclosures, exposing failure modes that single-turn, domain-agnostic evaluations cannot reveal.

大模型安全金融AI红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。