arXiv:2605.24210cs.LGstat.ML2026-05

揭示神经过程架构的表达能力层级,指导模型选型。

Characterizing the Representational Capacity of Neural Processes

  • 构建神经过程表达能力理论框架,分析四类主流架构
  • 发现CNPs到TNPs存在严格层级,每层支持更复杂上下文交互
  • 适合研究模型表达力或设计新神经过程架构的读者

神经过程能表示哪些函数?我们分析了条件神经过程(CNPs)、注意力神经过程(ANPs)、Transformer神经过程(TNPs)及其隐变量变体的表达能力。证明这些架构构成严格层级:仅依赖上下文分布有限个期望特征的函数可由CNPs表示;ANPs通过查询相关重加权实现核平滑,严格推广了CNPs;ConvCNPs与ANPs不可比较,分别对应平稳性与平移等变性;具有L层自注意力的TNPs可捕捉L跳上下文交互。对于隐变量神经过程,有限维隐变量虽支持一致采样,但无法绕过编码器限制;匹配高斯过程后验需隐维数随上下文规模增长。这些结果为基于任务结构选择架构提供了理论基础。

原文摘要 · Abstract (English)

What functions can Neural Processes represent? We analyze the representational capacity of popular NP architectures: Conditional Neural Processes (CNPs), Attentive Neural Processes (ANPs), Transformer Neural Processes (TNPs), and their latent variants. We prove these architectures form a strict hierarchy. CNP-representable functions are exactly those depending on finitely many expected features of the context distribution. ANPs strictly generalize CNPs via query-dependent reweighting, enabling kernel smoothers. ConvCNPs and ANPs are incomparable; each contains functions outside the other, separated by stationarity versus translation equivariance. TNPs with $L$ self-attention layers capture $L$-hop context interactions. For latent NPs, we show finite-dimensional latents provide coherent sampling but do not circumvent encoder limitations; matching GP posterior distributions requires latent dimension scaling with context size. These results provide a theoretical foundation for architecture selection based on task structure.

神经过程表达能力理论分析模型架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。