构建生成式AI评估科学,提升真实场景下的性能与安全可靠性。
Toward an Evaluation Science for Generative AI Systems
- 从交通、航空等领域借鉴安全评估经验,建立可落地的评估体系。
- 评估指标需贴近真实使用场景,且随技术演进持续优化。
- 适合关注AI安全性与可信评估的研究者与政策制定者阅读。
生成式AI系统在实际部署中的性能与安全性亟需前瞻性评估。然而当前评估体系存在不足:通用静态基准面临有效性挑战,临时性个案审计难以规模化。本文倡导发展生成式AI的评估科学。尽管生成式AI带来独特安全工程与测量难题,但可借鉴交通、航空航天及制药工程等领域的安全评估实践。我们总结三点关键启示:评估指标必须适用于真实世界表现,需迭代优化;必须建立专门的评估机构与行业规范。基于这些洞见,本文提出一条切实可行的路径,推动生成式AI评估向更严谨方向发展。
原文摘要 · Abstract (English)
There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluation ecosystem is insufficient: Commonly used static benchmarks face validity challenges, and ad hoc case-by-case audits rarely scale. In this piece, we advocate for maturing an evaluation science for generative AI systems. While generative AI creates unique challenges for system safety engineering and measurement science, the field can draw valuable insights from the development of safety evaluation practices in other fields, including transportation, aerospace, and pharmaceutical engineering. In particular, we present three key lessons: Evaluation metrics must be applicable to real-world performance, metrics must be iteratively refined, and evaluation institutions and norms must be established. Applying these insights, we outline a concrete path toward a more rigorous approach for evaluating generative AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。