arXiv:2504.20086cs.CLcs.AI2025-04中稿 · FAccT 2025被引 12

聚焦金融领域生成式AI风险,构建内容安全分类体系并验证现有防护措施失效。

Understanding and Mitigating Risks of Generative AI in Financial Services

  • 构建金融场景下生成式AI内容风险的分类体系
  • 红队测试显示现有技术防护措施无法识别多数风险内容
  • 适合金融AI监管、合规及安全团队参考

为负责任地开发生成式AI(GenAI)产品,明确可接受输入输出范围至关重要。何为“安全”响应仍存争议。学术研究多关注通用场景下的毒性、偏见与公平性评估,尤其在面向大众的对话应用中;但对专业化领域中的社会技术系统关注不足。而这些专业系统常面临严格的法律与监管审查。因此,需将产品特定考量纳入行业法规与企业治理要求。本文旨在突出金融服务业中生成式AI的内容安全问题,提出相应的风险分类体系。通过对比现有工作,分析各风险类别违规对利益相关方的影响,并利用红队测试数据评估现有开源技术防护方案的覆盖情况。结果表明,这些防护机制未能有效检测我们讨论的多数内容风险。

原文摘要 · Abstract (English)

To responsibly develop Generative AI (GenAI) products, it is critical to define the scope of acceptable inputs and outputs. What constitutes a "safe" response is an actively debated question. Academic work puts an outsized focus on evaluating models by themselves for general purpose aspects such as toxicity, bias, and fairness, especially in conversational applications being used by a broad audience. In contrast, less focus is put on considering sociotechnical systems in specialized domains. Yet, those specialized systems can be subject to extensive and well-understood legal and regulatory scrutiny. These product-specific considerations need to be set in industry-specific laws, regulations, and corporate governance requirements. In this paper, we aim to highlight AI content safety considerations specific to the financial services domain and outline an associated AI content risk taxonomy. We compare this taxonomy to existing work in this space and discuss implications of risk category violations on various stakeholders. We evaluate how existing open-source technical guardrail solutions cover this taxonomy by assessing them on data collected via red-teaming activities. Our results demonstrate that these guardrails fail to detect most of the content risks we discuss.

生成式AI金融风控内容安全模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。