探索如何让机器像科学家一样自主发现知识并解释发现。
Artificial Scientific Discovery
- 提出解释学习框架,让机器学会用自身语言解释新发现。
- 在猜谜游戏Zendo中成功实现自主科学发现与解释。
- 揭示大模型缺理解能力,需自建解释语言体系。
基于过去十年深度学习的爆发,本文从AlphaGo到ChatGPT,实证研究实现人工智能科学家所需的核心概念:机器自主生成原创研究并拓展人类知识的能力。研究始于Olivaw——一个类AlphaGo Zero的智能体,能从零发现奥赛罗棋知识,却无法有效传达。由此催生解释学习(EL)框架,形式化科学家向同行解释新现象的难题。有效应用该框架后,成功破解模拟科学探索的棋盘游戏Zendo。关键洞见在于:人工智能科学家必须发展自己对解释语言的诠释,而非依赖固定预设解释器。进一步审视现代多模态模型内部机制,提出一种解耦感知与解释的低成本方法,仅用少量多模态数据和无需额外训练即可构建类CLIP模型。最后讨论ChatGPT等模型仍欠缺成为真正科学探索者的能力,并引入Big-Bench符号解释任务,结果显示大型语言模型表现仅如随机猜测,而人类可完全解决。
原文摘要 · Abstract (English)
Rooted in the explosion of deep learning over the past decade, this thesis spans from AlphaGo to ChatGPT to empirically examine the fundamental concepts needed to realize the vision of an artificial scientist: a machine with the capacity to autonomously generate original research and contribute to the expansion of human knowledge. The investigation begins with Olivaw, an AlphaGo Zero-like agent that discovers Othello knowledge from scratch but is unable to communicate it. This realization leads to the development of the Explanatory Learning (EL) framework, a formalization of the problem faced by a scientist when trying to explain a new phenomenon to their peers. The effective EL prescriptions allow us to crack Zendo, a popular board game simulating the scientific endeavor. This success comes with a fundamental insight: an artificial scientist must develop its own interpretation of the language used to explain its findings, and not rely on a rigid existing interpreter. Questioning the very process of learning an interpreter, we turn our attention to the inner functioning of modern multimodal models. This culminates in a simple idea to build CLIP-like models where interpretation and perception are explicitly disentangled: a cost-effective approach that couples two unimodal models using little multimodal data and no further training. Finally, we discuss what ChatGPT and its siblings are still missing to become artificial scientists, and introduce the Big-Bench Symbol Interpretation Task, a benchmark about interpreting Zendo-like explanations that sees LLMs going no further than random chance while being instead fully solved by humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。