提出伦理测试框架,系统识别生成式AI内容的潜在危害
Ethics Testing: Proactive Identification of Generative AI System Harms

- 构建伦理测试方法,主动发现生成内容中的不当行为
- 通过五个案例验证该方法可有效检测代码与内容风险
- 适合关注AI安全与合规的研究者和开发者使用
生成式人工智能(GAI)系统凭借大语言模型(LLMs)在自动生成代码、图像等内容方面日益普及,但其生成内容的滥用可能引发严重后果。尽管确保生成内容质量至关重要,目前尚无系统性方法用于识别这些GAI系统产生的软件危害。本文提出“伦理测试”新概念,旨在系统化生成测试用例以发现生成内容中的伦理风险,如不道德行为或侵犯知识产权等问题。不同于现有的公平性测试,伦理测试聚焦于由不当行为引发的潜在危害。我们阐述了该方法面临的关键挑战,并通过五个案例研究展示了其在实际GAI系统中的应用效果。
原文摘要 · Abstract (English)
Generative Artificial Intelligence (GAI) systems that can automatically generate content in the form of source code or other contents (e.g., images) has seen increasing popularity due to the emergence of tools such as ChatGPT which rely on Large Language Models (LLMs). Misuse of the automatically generated content can incur serious consequences due to potential harms in the generated content. Despite the importance of ensuring the quality of automatically generated content, there is little to no approach that can systematically generate tests for identifying software harms in the content generated by these GAI systems. In this article, we introduce the novel concept of ethics testing which aims to systematically generate tests for identifying software harms. Different from existing testing methodologies (e.g., fairness testing that aims to identifying software discrimination), ethics testing aims to systematically detect software harms that could be induced due to unethical behavior (e.g., harmful behavior or behavior that violates intellectual property rights) in automatically generated content. We introduced the concept of ethics testing, discussed the challenges therewithin, and conducted five case studies to show how ethics testing can be performed for generative AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。