构建金融安全评测基准,检验大模型在真实场景下的合规拒答能力
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
- 基于真实金融犯罪案例设计双语红队测试集
- 发现通用与金融专用模型均存在合规绕过漏洞
- 中文场景下模型更易被诱导,提示防御效果有限
大型语言模型在金融场景中应用日益广泛,但可能生成有害输出,如协助非法活动或违背伦理行为,带来严重合规风险。为系统评估金融场景下大模型的安全性,我们提出FinSafetyBench,一个涵盖14个子类别的双语(中英文)红队评测基准,用于测试模型对违反金融合规请求的拒绝能力。该基准基于真实金融犯罪案例与伦理标准,覆盖金融犯罪与道德违规。在三种典型攻击设置下对通用与金融专用大模型进行大量实验,发现对抗性提示可绕过合规防护机制。进一步分析显示,中文场景下模型更易受攻击,且提示级防御难以应对复杂或隐含操纵策略。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal activities or unethical behavior, posing serious compliance risks. To systematically evaluate LLM safety in finance, we propose FinSafetyBench, a bilingual (English-Chinese) red-teaming benchmark designed to test an LLM's refusal of requests that violate financial compliance. Grounded in real-world financial crime cases and ethics standards, the benchmark comprises 14 subcategories spanning financial crimes and ethical violations. Through extensive experiments on general-purpose and finance-specialized LLMs under three representative attack settings, we identify critical vulnerabilities that allow adversarial prompts to bypass compliance safeguards. Further analysis reveals stronger susceptibility in Chinese contexts and highlights the limitations of prompt-level defenses against sophisticated or implicit manipulation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。