构建模型科学框架,让AI更可验、可解释、可控制。
Model Science: getting serious about verification, explanation and control of AI systems
- 以模型为中心,建立验证、解释、控制与交互四大支柱。
- 强调上下文感知的严格评估与对模型内部机制的探索。
- 适合关注AI可信性与安全性的研究人员和开发者。
基础模型日益普及,推动数据科学向模型科学范式转变。与以数据为中心的方法不同,模型科学将训练好的模型置于分析核心,旨在跨多种应用场景中与模型交互、验证其行为、解释其机制并实现有效控制。本文提出一个名为模型科学的新学科概念框架,包含四个关键支柱:验证(要求严格且上下文敏感的评估协议)、解释(涵盖探索模型内部运作的各种方法)、控制(整合对齐技术以引导模型行为)以及接口(开发交互式可视化工具,提升人类校准与决策能力)。该框架旨在指导可信、安全且与人类对齐的AI系统发展。
原文摘要 · Abstract (English)
The growing adoption of foundation models calls for a paradigm shift from Data Science to Model Science. Unlike data-centric approaches, Model Science places the trained model at the core of analysis, aiming to interact, verify, explain, and control its behavior across diverse operational contexts. This paper introduces a conceptual framework for a new discipline called Model Science, along with the proposal for its four key pillars: Verification, which requires strict, context-aware evaluation protocols; Explanation, which is understood as various approaches to explore of internal model operations; Control, which integrates alignment techniques to steer model behavior; and Interface, which develops interactive and visual explanation tools to improve human calibration and decision-making. The proposed framework aims to guide the development of credible, safe, and human-aligned AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。