医学机构用红队测试发现生成式AI可能泄露版权内容,已部署防护措施。
Red Teaming for Generative AI, Report on a Copyright-Focused Exercise Completed in an Academic Medical Center
- 组织42人开展红队测试,模拟提取四类文本中的受版权保护内容。
- 成功复现书籍致谢词和近似原文段落,但新闻与临床笔记未被复制。
- 测试结果推动上线专门防版权违规的元提示,适用于医疗科研机构。
在学术医疗环境中部署生成式人工智能引发版权合规担忧。丹娜-法伯癌症研究所推出GPT4DFCI,基于OpenAI模型并获企业级使用许可,用于研究与运营。鉴于该工具在机构内广泛使用、研究使命驱动以及需遵循Azure OpenAI服务客户版权承诺的共担责任模式,我们实施了严格的版权合规测试。2024年11月,来自学术界、产业界和政府机构的42名参与者组成四支团队,在文学作品、新闻文章、科学出版物及受限临床记录四个领域尝试从GPT4DFCI中提取受版权保护内容。团队成功提取出原文献致谢词和近乎完整的段落,尽管通过越狱攻击仍未能提取新闻内容;科学文献仅返回高层摘要;临床记录测试显示隐私保护机制有效。结果显示,文学内容可被复现,表明训练数据中可能存在版权材料,需在推理阶段增加过滤机制。不同内容类型的成功率差异说明系统保护策略存在差异。本次事件促成在GPT4DFCI中部署专门的版权防护元提示,自2025年1月起已投入生产。结论:系统性红队测试揭示了生成式人工智能在版权合规方面的具体漏洞,并催生可落地的缓解方案。学术医疗机构应建立持续测试机制以保障法律与伦理合规。
原文摘要 · Abstract (English)
Background: Generative artificial intelligence (AI) deployment in academic medical settings raises copyright compliance concerns. Dana-Farber Cancer Institute implemented GPT4DFCI, an internal generative AI tool utilizing OpenAI models, that is approved for enterprise use in research and operations. Given (1) the exceptionally broad adoption of the tool in our organization, (2) our research mission, and (3) the shared responsibility model required to benefit from Customer Copyright Commitment in Azure OpenAI Service products, we deemed rigorous copyright compliance testing necessary. Case Description: We conducted a structured red teaming exercise in Nov. 2024, with 42 participants from academic, industry, and government institutions. Four teams attempted to extract copyrighted content from GPT4DFCI across four domains: literary works, news articles, scientific publications, and access-restricted clinical notes. Teams successfully extracted verbatim book dedications and near-exact passages through various strategies. News article extraction failed despite jailbreak attempts. Scientific article reproduction yielded only high-level summaries. Clinical note testing revealed appropriate privacy safeguards. Discussion: The successful extraction of literary content indicates potential copyrighted material presence in training data, necessitating inference-time filtering. Differential success rates across content types suggest varying protective mechanisms. The event led to implementation of a copyright-specific meta-prompt in GPT4DFCI; this mitigation has been in production since Jan. 2025. Conclusion: Systematic red teaming revealed specific vulnerabilities in generative AI copyright compliance, leading to concrete mitigation strategies. Academic medical institutions deploying generative AI should implement continuous testing protocols to ensure legal and ethical compliance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。