大模型通过上下文学习,能精准捕捉隐马尔可夫模型生成的序列规律。
Pre-trained Large Language Models Learn Hidden Markov Models In-context
- 利用上下文学习让大模型从示例中推断隐马尔可夫结构
- 在合成数据上预测准确率接近理论最优水平
- 为科研人员提供探测复杂数据隐藏结构的新工具
隐马尔可夫模型(HMM)是建模具有潜在马尔可夫结构的序列数据的基础工具,但将其拟合到真实数据仍面临计算挑战。本文发现,预训练大语言模型(LLMs)可通过上下文学习(ICL)有效建模由HMM生成的数据——即从提示中的示例推断模式。在多种合成HMM数据上,LLMs的预测准确率接近理论最优。我们揭示了受HMM特性影响的新规模效应,并提出理论猜想解释这些现象。同时,为科学家提供了使用ICL诊断复杂数据的实用指南。在真实世界动物决策任务中,ICL性能媲美人工设计的专业模型。据我们所知,这是首次证明ICL可学习并预测HMM生成序列,深化了对大模型上下文学习的理解,并确立其作为揭示复杂科学数据隐藏结构的强大工具的潜力。
原文摘要 · Abstract (English)
Hidden Markov Models (HMMs) are foundational tools for modeling sequential data with latent Markovian structure, yet fitting them to real-world data remains computationally challenging. In this work, we show that pre-trained large language models (LLMs) can effectively model data generated by HMMs via in-context learning (ICL)$\unicode{x2013}$their ability to infer patterns from examples within a prompt. On a diverse set of synthetic HMMs, LLMs achieve predictive accuracy approaching the theoretical optimum. We uncover novel scaling trends influenced by HMM properties, and offer theoretical conjectures for these empirical observations. We also provide practical guidelines for scientists on using ICL as a diagnostic tool for complex data. On real-world animal decision-making tasks, ICL achieves competitive performance with models designed by human experts. To our knowledge, this is the first demonstration that ICL can learn and predict HMM-generated sequences$\unicode{x2013}$an advance that deepens our understanding of in-context learning in LLMs and establishes its potential as a powerful tool for uncovering hidden structure in complex scientific data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。