arXiv:2508.10108cs.AIcs.CL2025-08被引 1

十支高校团队竞赛构建安全编码AI,对抗测试验证模型可靠性

Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development

  • 通过对抗性比赛让AI助手与红队机器人多轮对话,测试安全性
  • 五支团队开发出能检测多轮越狱攻击的新型安全对齐方法
  • 适合关注AI代码生成安全、模型对抗防御的研究者和开发者

用于软件开发的AI系统日益普及,但其安全性仍面临重大挑战。为此,亚马逊发起了亚马逊新星AI挑战赛的可信AI赛道,汇集全球10所高校团队,推动安全AI的发展。其中5支队伍专注于开发自动化红队机器人,另5支则致力于构建安全的AI编程助手。挑战赛设置对抗性锦标赛,红队与参赛AI助手进行多轮对话,测试其安全对齐能力;同时提供高质量标注数据流,支持团队迭代优化。在比赛中,各团队提出创新方法,涵盖基于推理的安全对齐、鲁棒的模型防护机制、多轮越狱攻击检测及大语言模型高效探测技术。为支持研究,亚马逊新星团队投入大量科研与工程资源,包括从零构建基准编码模型、开发赛事调度服务与评估框架。本文总结了高校团队与亚马逊团队在提升软件开发AI安全性方面取得的关键进展,展示了协同攻关对推动行业安全标准的意义。

原文摘要 · Abstract (English)

AI systems for software development are rapidly gaining prominence, yet significant challenges remain in ensuring their safety. To address this, Amazon launched the Trusted AI track of the Amazon Nova AI Challenge, a global competition among 10 university teams to drive advances in secure AI. In the challenge, five teams focus on developing automated red teaming bots, while the other five create safe AI assistants. This challenge provides teams with a unique platform to evaluate automated red-teaming and safety alignment methods through head-to-head adversarial tournaments where red teams have multi-turn conversations with the competing AI coding assistants to test their safety alignment. Along with this, the challenge provides teams with a feed of high quality annotated data to fuel iterative improvement. Throughout the challenge, teams developed state-of-the-art techniques, introducing novel approaches in reasoning-based safety alignment, robust model guardrails, multi-turn jail-breaking, and efficient probing of large language models (LLMs). To support these efforts, the Amazon Nova AI Challenge team made substantial scientific and engineering investments, including building a custom baseline coding specialist model for the challenge from scratch, developing a tournament orchestration service, and creating an evaluation harness. This paper outlines the advancements made by university teams and the Amazon Nova AI Challenge team in addressing the safety challenges of AI for software development, highlighting this collaborative effort to raise the bar for AI safety.

AI安全代码生成对抗测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。