开放AI如何更安全?这场会议给出了五条关键研究方向。
A Different Approach to AI Safety: Proceedings from the Columbia Convening on Openness in Artificial Intelligence and AI Safety
- 通过多方协作,提出安全与开源AI交叉的研究议程。
- 发现多模态/多语言基准缺失,智能体防御能力不足。
- 适合关注AI治理、安全开发与政策制定的读者。
本文报告了2024年11月19日于旧金山举行的哥伦比亚人工智能开放性与安全研讨会及其为期六周的筹备项目成果。来自学术界、产业界、民间社会和政府的四十五位以上研究人员、工程师和政策领袖参与,采用参与式、解决方案导向的方法,形成三方面成果:(一)安全与开源AI交叉领域的研究议程;(二)覆盖AI开发全链路的现有与亟需技术干预及开源工具图谱;(三)内容安全过滤生态的映射及未来研究路线图。研究发现,透明权重、可互操作工具与公共治理形式的开放性,可通过独立审查、去中心化缓解和多元文化监督增强系统安全。但仍有显著短板:多模态与多语言基准稀缺,智能体系统对提示注入与组合攻击防御有限,受AI伤害影响最深群体的参与机制不足。论文最终提出五项优先研究方向:强调参与式输入、未来可扩展的内容过滤器、全生态安全基础设施、严格的智能体防护机制以及扩展的伤害分类体系。这些建议已为2025年2月法国人工智能行动峰会提供依据,并为建立开放、多元、负责的AI安全学科奠定基础。
原文摘要 · Abstract (English)
The rapid rise of open-weight and open-source foundation models is intensifying the obligation and reshaping the opportunity to make AI systems safe. This paper reports outcomes from the Columbia Convening on AI Openness and Safety (San Francisco, 19 Nov 2024) and its six-week preparatory programme involving more than forty-five researchers, engineers, and policy leaders from academia, industry, civil society, and government. Using a participatory, solutions-oriented process, the working groups produced (i) a research agenda at the intersection of safety and open source AI; (ii) a mapping of existing and needed technical interventions and open source tools to safely and responsibly deploy open foundation models across the AI development workflow; and (iii) a mapping of the content safety filter ecosystem with a proposed roadmap for future research and development. We find that openness -- understood as transparent weights, interoperable tooling, and public governance -- can enhance safety by enabling independent scrutiny, decentralized mitigation, and culturally plural oversight. However, significant gaps persist: scarce multimodal and multilingual benchmarks, limited defenses against prompt-injection and compositional attacks in agentic systems, and insufficient participatory mechanisms for communities most affected by AI harms. The paper concludes with a roadmap of five priority research directions, emphasizing participatory inputs, future-proof content filters, ecosystem-wide safety infrastructure, rigorous agentic safeguards, and expanded harm taxonomies. These recommendations informed the February 2025 French AI Action Summit and lay groundwork for an open, plural, and accountable AI safety discipline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。