arXiv:2606.01189cs.AI2026-06被引 2

AI模型需从比拼性能转向深入理解其工作原理。

The Case for Model Science: Verify, Explore, Steer, Refine

论文配图:The Case for Model Science: Verify, Explore, Steer, Refine
图 1 · 摘自论文原文
  • 提出验证、探索、引导、优化四维分析框架。
  • 强调单个模型深度剖析比群体研究更能发现隐藏问题。
  • 适合研究者与工程师构建可积累的模型分析体系。

我们主张,人工智能界已具备条件超越基准测试,将零散的模型分析整合为系统性学科——模型科学。当前复杂AI模型服务数十亿用户,但对其工作机制的理解远落后于部署能力。数十年基于基准的研宄虽带来显著进展,如全面排行榜和多样化指标,但无法解释模型成功或失败的原因,也遗漏了幻觉、捷径等关键失效模式。借鉴认知科学、神经科学、医学与农业的经验:理解复杂系统需多层级分析;单个案例深究能揭示群体研究忽略的细节;专业训练须与研究并行;共享基础设施推动持续进步。据此提出模型科学三大基础:一、构建验证、探索、引导、优化四类功能视角,覆盖模型行为的不同问题;二、建立包含数据集、模型与发现的目录化知识库;三、重视单个模型实例的深度分析,因其可能揭示群体研究无法察觉的特性。

原文摘要 · Abstract (English)

We argue that the AI community is now ready to move beyond benchmarking and consolidate scattered efforts in model analysis into a systematic discipline, a direction we term Model Science. Complex AI models now serve billions of users, yet our understanding of how they work lags far behind our ability to deploy them. Decades of benchmark-driven research have delivered remarkable progress: extensive leaderboards, a wide range of performance metrics, tracking capability gains across diverse tasks; yet this success has also revealed the limits of benchmarks as they tell us whether models perform but not why they succeed or fail, they miss critical failure modes, such as hallucinations or shortcuts. Precedents from established sciences point the way forward: cognitive science shows that understanding complex systems requires complementary levels of analysis; neuroscience demonstrates that deep study of single cases reveals what population studies miss; medicine teaches that specialised training must develop alongside research practice; and agriculture models how shared infrastructure and principles enable cumulative progress. These lessons inform three foundations for Model Science. First, we propose to consolidate research around four functional perspectives: Verify, Explore, Steer, and Refine that address complementary questions about model behaviour. Second, we discuss the required infrastructure for cumulative knowledge: catalogues of datasets, models and findings. Third, we highlight the need for deep analysis of individual model instances, not just model families, because single cases can reveal what population studies miss.

模型分析认知科学系统性研究知识积累

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。