构建SAIF框架,系统评估生成式AI在公共部门的风险。
SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector
- 提出四阶段风险生成框架,系统化设计测试场景
- 涵盖多模态能力与新型越狱方法,覆盖广泛风险类型
- 适合政策制定者与AI安全研究人员参考
生成式AI在公共部门的快速应用,涵盖自动化公共服务、福利服务和移民流程等多样场景,展现了变革潜力,也凸显了全面风险评估的紧迫性。尽管其广泛应用,但针对公共部门中人工智能系统风险的评估仍不充分。基于政府政策与企业指南中的现有AI风险分类体系,本研究深入分析生成式AI在公共部门带来的关键风险,并扩展范围以涵盖其多模态特性。为此,我们提出系统性数据生成框架(SAIF),包含四个关键阶段:风险拆解、场景设计、越狱方法应用与提示类型探索。SAIF确保提示数据的系统化与一致性生成,支持全面评估并为风险缓解提供基础。此外,该框架可适应新兴越狱方法与不断演进的提示类型,从而有效应对未知风险场景。我们认为本研究对推动生成式AI在公共部门的安全与负责任融合具有重要意义。
原文摘要 · Abstract (English)
The rapid adoption of generative AI in the public sector, encompassing diverse applications ranging from automated public assistance to welfare services and immigration processes, highlights its transformative potential while underscoring the pressing need for thorough risk assessments. Despite its growing presence, evaluations of risks associated with AI-driven systems in the public sector remain insufficiently explored. Building upon an established taxonomy of AI risks derived from diverse government policies and corporate guidelines, we investigate the critical risks posed by generative AI in the public sector while extending the scope to account for its multimodal capabilities. In addition, we propose a Systematic dAta generatIon Framework for evaluating the risks of generative AI (SAIF). SAIF involves four key stages: breaking down risks, designing scenarios, applying jailbreak methods, and exploring prompt types. It ensures the systematic and consistent generation of prompt data, facilitating a comprehensive evaluation while providing a solid foundation for mitigating the risks. Furthermore, SAIF is designed to accommodate emerging jailbreak methods and evolving prompt types, thereby enabling effective responses to unforeseen risk scenarios. We believe that this study can play a crucial role in fostering the safe and responsible integration of generative AI into the public sector.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。