用因子化遗憾度量变量交互对推理的影响,提升组合泛化能力。
Factorization Regret mediates compositional generalization in latent space
- 引入因子化遗憾,量化潜变量交互对任务性能的贡献。
- RCC架构在稀疏奖励下实现潜变量交互学习,支持新组合泛化。
- 揭示了置信度与准确率脱钩的理论失效模式,适合研究智能体泛化。
当所有相关变量已知时,是否存在泛化障碍?我们研究了如何泛化到任务相关潜变量的新组合。为此,我们构建了认知网格世界(Cognitive Gridworld),一个静态部分可观测马尔可夫决策过程(POMDP),其中观测由多个具有参数化交互的潜变量共同生成。该设置允许定义因子化遗憾:一个信息论量,用于衡量潜变量交互对任务性能的贡献。通过该度量,我们分析了显式提供交互信息的循环神经网络(RNNs),发现因子化遗憾能解释回声状态网络与全训练网络之间的准确率差距。此外,我们的分析揭示了一种理论预测的失效模式——置信度与准确率脱钩。随后,我们处理更难的情形:交互关系需由嵌入模型自行学习。在同时推断变量值和估计交互关系的变分推断问题中,我们提出表示分类链(RCCs)——一种显式分离变量推断与参数估计的新架构。RCCs可在仅部分目标变量有稀疏奖励的情况下学习交互关系。最终,我们在离线学习中验证了RCCs在新组合泛化上的有效性。综上,我们提出了一个理论坚实的研究框架,用于目标导向智能的开发与评估。
原文摘要 · Abstract (English)
Are there still barriers to generalization once all of the relevant variables are known? We consider the challenge of generalizing to a novel combination of task-relevant latent variables. To explore this framework, we develop the Cognitive Gridworld, a stationary Partially Observable Markov Decision Process (POMDP) in which observations are generated jointly by multiple latent variables with parametric interactions. This setting allows us to describe Factorization Regret: an information-theoretic quantity that measures the contribution of latent variable interactions to task performance. Using this metric, we first analyze Recurrent Neural Networks (RNNs) that are explicitly provided with the interactions and find that Factorization Regret explains the accuracy gap between Echo State and Fully Trained networks. Additionally, our analysis uncovers a theoretically predicted failure mode, where confidence becomes decoupled from accuracy. These results suggest that utilizing the interactions between relevant variables is a non-trivial capability. We then address a harder regime where the interactions themselves must be learned by an embedding model. Estimating how variables interact while simultaneously inferring their values is a variational inference problem. To explicitly disentangle variable inference from parameter estimation, we develop Representation Classification Chains (RCCs), a novel architecture for learning how latent variables interact. RCCs are capable of learning from sparse and partial reward, provided only for a subset of goal variables. Finally, we demonstrate the usefulness of RCCs in enabling generalization to novel combinations of latent variables through offline learning. In summary, we present a theoretically grounded setting for research, development and evaluation of goal-directed general intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。