arXiv:2606.06533cs.AIcs.CL2026-06中稿 · ICML被引 1

研究训练过程本身,才能真正理解AI行为的成因。

Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics

  • 把模型看作动态演化过程,而非静态结果。
  • 需从早期信号预测模型能力、偏见等关键属性。
  • 适合关注AI可解释性与可靠性研究者阅读。

AI 的科学理解应超越对训练后模型的分析,转而研究塑造模型行为的训练动态。当前研究多将模型视为固定产物,忽视其在数据、目标、架构和优化过程中的演化本质。本文主张,真正的科学需能基于早期训练信号预测模型的能力、偏见、鲁棒性与安全性等表现,并实现对训练轨迹的干预与设计。虽然损失的缩放定律已实现预测,但将此成功扩展至能力、偏见等复杂行为仍是挑战。文章结合科学哲学,探讨机制可解释性、公平性、记忆现象与简单性偏差等方面的进展,提出具体开放问题。

原文摘要 · Abstract (English)

What would it mean to have a scientific understanding of AI? Models are not static objects: they are snapshots of time-evolving processes shaped by data, objectives, architectures, and optimization dynamics. Yet much of AI research treats models as fixed artifacts, analyzing behaviors after training rather than asking why they emerge. This position paper argues that a science of AI must move beyond post-hoc fixes and study the training dynamics that produce model behavior. Such a science should support progressively stronger forms of understanding: predicting outcomes from early training signals, intervening when trajectories go wrong, and ultimately designing training procedures that more reliably produce desired properties. Scaling laws have made prediction routine for loss; the challenge is extending this success to capabilities, biases, robustness, and safety-relevant behaviors. We articulate requirements for such theories grounded in the history and philosophy of science, examine progress in mechanistic interpretability, fairness, memorization, and simplicity bias, and identify concrete open problems.

训练动态可解释性模型行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。