改进Transformer架构,让模型能零样本预测动态系统行为
Enhanced Transformer architecture for in-context learning of dynamical systems
- 用概率框架建模,提升对不确定性的表达能力
- 支持不连续上下文,增强实际场景适应性
- 引入递归分块处理长序列,提升可扩展性
近期由部分作者提出,在上下文识别范式中,可通过合成数据离线训练一个元模型,以描述一类系统的整体行为。模型训练完成后,将真实系统生成的输入/输出序列(上下文)输入,即可实现零样本学习预测其行为。本文通过三项关键改进增强原元建模框架:在概率框架下定义学习任务;支持非连续上下文与查询窗口;采用递归分块有效处理长上下文序列。数值实验以魏纳-霍克斯坦系统类为例,验证了模型性能提升与可扩展性。
原文摘要 · Abstract (English)
Recently introduced by some of the authors, the in-context identification paradigm aims at estimating, offline and based on synthetic data, a meta-model that describes the behavior of a whole class of systems. Once trained, this meta-model is fed with an observed input/output sequence (context) generated by a real system to predict its behavior in a zero-shot learning fashion. In this paper, we enhance the original meta-modeling framework through three key innovations: by formulating the learning task within a probabilistic framework; by managing non-contiguous context and query windows; and by adopting recurrent patching to effectively handle long context sequences. The efficacy of these modifications is demonstrated through a numerical example focusing on the Wiener-Hammerstein system class, highlighting the model's enhanced performance and scalability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。