arXiv:2509.10546cs.CLcs.AI2025-09ACL被引 4

提出可控制的多轮隐蔽风险攻击框架,提升金融大模型安全测试效果

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain

  • 通过迭代优化生成表面合规但诱导违规响应的多轮提示
  • 在9个主流模型上实现95%的平均攻击成功率
  • 适用于金融领域大模型安全评估与对抗训练

大语言模型在金融领域的应用日益广泛,但不当行为可能引发严重监管风险。现有红队测试多关注明显有害内容,忽视表面合法却诱导违规回复的隐蔽攻击。本文提出可控黑盒多轮隐蔽风险红队框架CoRT,通过迭代生成逐步隐藏表层风险的多轮提示,并引入风险隐蔽度控制器(RCC)实时预测风险隐蔽分数以引导后续生成风格。构建了包含522条指令的金融风险专用基准数据集FinRisk-Bench,覆盖六大金融风险类别。在九个主流大模型上的实验表明,CoRT(RCA)平均攻击成功率达93.19%,加入RCC后进一步提升至95.00%。代码与数据集已开源。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in finance, where unsafe behavior can lead to serious regulatory risks. However, most red-teaming research focuses on overtly harmful content and overlooks attacks that appear legitimate on the surface yet induce regulatory-violating responses. We address this gap by introducing a controllable black-box multi-turn risk-concealed red-teaming framework (CoRT) that progressively conceals surface-level risk while exploiting regulatory-violating behaviors. CoRT contains two key components: (i) a Risk Concealment Attacker (RCA) that generates multi-turn prompts via iterative refinement, and (ii) a Risk Concealment Controller (RCC) that predicts a turn-level Risk Concealment Score (RCS) to steer RCA's follow-up style. We also built a domain-specific benchmark, FinRisk-Bench, with 522 instructions spanning six financial risk categories. Experiments on nine widely used LLMs show that CoRT (RCA) achieves 93.19% average attack success rate (ASR), and CoRT (RCA+RCC) further improves the average ASR to 95.00%. Our code and FinRisk-Bench are available at https://github.com/gcheng128/CoRT.

大模型安全红队测试金融AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。