提出新型结构化动态模型,验证其在特定任务中的优势
Explicit Interaction Architectures for Dynamical Learning: A Controlled Study of Structural Inductive Bias
- 基于波模型设计有向状态交互单元,避免代数环
- 单层结构在非线性识别任务上误差低至2.76e-4
- 结果表明结构先验效果依赖任务,不具普适性
我们研究了一种结构优先的动力学学习方法,其中状态交互的组织形式被显式规定,而非完全依赖通用循环参数化。提出由有序局部状态调制变换构成的因果循环单元,其设计受波传播模型启发,但未施加散射、无源或能量守恒约束。鉴于固定循环动力学、设计的储备池拓扑、仅读出学习及循环深度已广为人知,本研究聚焦于:在受控计算条件下,所提交互结构是否提供有效归纳偏置?对比了单层、双层结构模型与通用回声状态网络(ESN),三者均含12个循环状态和相同线性岭回归读出。所有模型在独立校准数据上使用相同随机搜索预算,之后冻结超参。在自定义非线性识别任务中,单层结构平均验证归一化均方误差为2.76×10⁻⁴,优于双层模型的3.19×10⁻⁴和ESN的3.94×10⁻⁴;而在NARMA10任务中,结果反转:ESN为0.312,单层和双层分别为0.348和0.357。表明该结构可竞争且有利,但非普遍最优;且在相同状态维度下,循环深度并未系统提升性能。结果支持任务依赖的结构归纳偏置观点,并将该架构定位为更强波与系统理论构造的可控前导。
原文摘要 · Abstract (English)
We investigate a structure-first approach to dynamical learning in which the organization of stateful interactions is prescribed explicitly rather than left entirely to a generic recurrent parameterization. We introduce causal recurrent units built from an ordered sequence of local, state-modulated transformations. The construction is motivated by wave-based interaction models, but the units studied here do not impose scattering, passivity, or energy-balance constraints. Because fixed recurrent dynamics, designed reservoir topologies, readout-only learning, and recurrent depth are already well established, the empirical question is deliberately narrower: does the proposed interaction organization provide a useful inductive bias under controlled computational conditions? We compare a one-layer structured model, a two-layer structured model, and a generic echo-state network (ESN), all with 12 recurrent states and the same strictly linear ridge readout. Each model family receives the same random-search budget on calibration data that are disjoint from the final test data, after which the selected hyperparameters are frozen. On a custom nonlinear identification task, the one-layer structured model attains a mean validation NMSE of 2.76 x 10^-4, compared with 3.19 x 10^-4 for the two-layer model and 3.94 x 10^-4 for the ESN. On NARMA10 the ordering reverses: the ESN attains 0.312, compared with 0.348 and 0.357 for the one- and two-layer structured models. Thus, the proposed organization can be competitive and advantageous on one task, but it is not universally superior; moreover, recurrent depth does not provide a systematic benefit under matched state dimension. The results support a task-dependent interpretation of structural inductive bias and position the present architecture as a controlled precursor to stronger wave- and system-theoretic constructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。