arXiv:2606.19887cs.CRcs.AI2026-06被引 1

为金融大模型量身打造安全评估框架,精准识别欺诈与合规风险。

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

论文配图:FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming
图 1 · 摘自论文原文
  • 基于金融专家设计的两级威胁分类体系,关联国际监管标准。
  • 将真实金融文档转为高保真对抗提示,减少28到12例关键漏检。
  • 已用于韩国金融安全研究所实证评估,限流开放确保安全使用。

现有安全评测基准多聚焦通用对抗场景,忽视金融领域特有风险。金融大模型可能引发监管违规、助长欺诈、削弱系统信任,亟需针对性评估。我们提出FinRED——一个由金融专家指导的红队测试框架,通过创新的两级分类体系,将全球标准(如FATF、EU DORA)映射至从规避监管到复杂欺诈的各类威胁,并构建可扩展的流水线,依据专家定义的模式将真实金融文档转化为情境丰富的红队行为提示(种子)。经专家严格验证,提示具备现实合理性,可支撑有效的大模型安全评估。同时提供经专家验证的金融专用评分标准,超越简单免责声明检查,更贴近人类专家判断,将关键误判率从28降至12。该框架符合国际信息安全管理标准(如ISO/IEC 27001),已在韩国金融安全研究所(FSI)生成式AI监管沙箱中投入实际应用。为防范双重用途风险,数据集、生成流程、提示模板与评估框架仅对合格研究者开放,详见https://github.com/selectstar-ai/FinRED-paper 和 https://huggingface.co/datasets/datumo/FinRED。

原文摘要 · Abstract (English)

Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance violations, fraud facilitation, and systemic trust erosion that require targeted evaluation. We introduce FinRED, an expert-guided red-teaming framework for financial LLM safety evaluation developed with financial experts. FinRED uses a novel two-level taxonomy mapping global standards (e.g., FATF and EU DORA) to threats ranging from regulatory evasion to complex fraud, integrated with a scalable pipeline that converts real financial documents into context-rich red-teaming Behavioral Prompts (seeds) through an expert-defined schema. Rigorous expert validation confirms seed plausibility and realism for meaningful LLM safety evaluation. We also provide an expert-validated, finance-specific rubric that goes beyond disclaimer checks, aligns more closely with human experts than static one-size-fits-all rubrics, and reduces critical false negatives from 28 to 12. Aligned with internationally adopted risk-management and information-security standards (e.g., ISO/IEC 27001), FinRED is deployed in South Korea's Financial Security Institute (FSI) regulatory sandbox for generative AI security evaluation in real financial services. To mitigate dual-use risks, the dataset, generation pipeline, prompt template, and evaluation framework are gated for qualified researchers at https://github.com/selectstar-ai/FinRED-paper and https://huggingface.co/datasets/datumo/FinRED.

金融AI红队测试安全评估大模型治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。