UNICON模型实现跨学科数值智能,无需重训即可达专家水平。
A foundation model of numerical intelligence with cross-disciplinary generalization

- 用图结构示例建模,从数值上下文中推断通用规律。
- 在未训练过的学科上表现接近甚至超越专业模型。
- 适合需要跨领域泛化能力的研究者与开发者。
智力通常被理解为获取和应用知识、适应陌生情境并解决新问题的能力。大型语言模型通过从文本上下文中推断任务相关知识并应用于新任务来体现这一能力。然而,智力并不局限于语言。对于科学与社会系统,我们需要能从数值上下文中获取和应用知识的模型——我们称之为数值智能。本文提出统一上下文操作网络(UNICON),一种在跨学科中展现数值智能的基础模型。利用系统内图结构示例作为上下文,UNICON 推断其共享的预测关系,并将其应用于同一系统的查询。在多个科学与社会系统中,包括训练时未涉及的学科,同一模型在不重新训练的情况下接近甚至达到专家性能。将 UNICON 与语言模型代理结合进行上下文集成学习(CEL),进一步提升表现,使其在训练中未见的学科上超越现有最先进专业模型。此外,训练语料多样性有助于提升对未见学科的泛化能力。这些结果确立了 UNICON 作为数值智能基础模型的地位,为其成为更广泛人工智能生态系统的构建模块奠定了基础。
原文摘要 · Abstract (English)
Intelligence is commonly understood as the ability to acquire and apply knowledge, adapt to unfamiliar situations and solve new problems. Large language models exhibit this capacity by inferring task-relevant knowledge from textual context and applying it to new tasks. Yet intelligence need not be confined to language. For scientific and social systems, we need models that acquire and apply knowledge from numerical context-an ability we call numerical intelligence. Here we introduce UNified In-Context Operator Networks (UNICON), a foundation model that exhibits numerical intelligence across disciplines. Using graph-based examples from a system as context, UNICON infers the predictive relation shared across them and applies it to queries from the same system. Across scientific and social systems, including those from disciplines absent from training, the same model approaches specialist performance without retraining. Combining UNICON with language-model agents to perform contextual ensemble learning (CEL) yields further gains, enabling it to surpass state-of-the-art specialists in a discipline unseen during training. We further show that training-corpus diversity improves generalization to unseen disciplines. Together, these results establish UNICON as a foundation model of numerical intelligence and position it as a building block for a broader ecosystem of artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。