arXiv:2501.14346cs.LGcs.AI2025-01中稿 · the ACML conferenc…

HorNets同时处理离散与连续表格数据,高效学习小样本高维数据。

HorNets: Learning from Discrete and Continuous Signals with Routing Neural Networks

  • 用自定义路由机制动态选择网络路径,依据输入特征基数决定计算模式。
  • 在14个生物医学数据集上达到顶尖分类性能,能准确提取逻辑规则(含噪声XNOR)。
  • 适合小样本、高维、混合类型表格数据的场景,开源可复现。

构建适用于同时学习连续与离散表格数据的神经网络架构是一项挑战性研究。当前高维表格数据集通常实例数量较少,需要数据高效的训练方式。我们提出HorNets(Horn Networks),在合成数据和真实世界的小样本表格数据集上表现卓越。HorNets基于截断多项式类激活函数,引入定制的离散-连续路由机制,根据输入特征的基数决定网络中哪部分进行优化。通过显式建模特征组合空间的部分或以类似线性注意力的方式整合整个空间,HorNets可无监督地动态选择最适合当前数据的运行模式。该架构是少数能稳定恢复逻辑命题(包括含噪XNOR)并实现14个真实生物医学高维数据集上顶尖分类性能的方法之一。HorNets已开源,附带用于生成类别基准的合成数据生成器。

原文摘要 · Abstract (English)

Construction of neural network architectures suitable for learning from both continuous and discrete tabular data is a challenging research endeavor. Contemporary high-dimensional tabular data sets are often characterized by a relatively small instance count, requiring data-efficient learning. We propose HorNets (Horn Networks), a neural network architecture with state-of-the-art performance on synthetic and real-life data sets from scarce-data tabular domains. HorNets are based on a clipped polynomial-like activation function, extended by a custom discrete-continuous routing mechanism that decides which part of the neural network to optimize based on the input's cardinality. By explicitly modeling parts of the feature combination space or combining whole space in a linear attention-like manner, HorNets dynamically decide which mode of operation is the most suitable for a given piece of data with no explicit supervision. This architecture is one of the few approaches that reliably retrieves logical clauses (including noisy XNOR) and achieves state-of-the-art classification performance on 14 real-life biomedical high-dimensional data sets. HorNets are made freely available under a permissive license alongside a synthetic generator of categorical benchmarks.

表格数据小样本学习路由网络生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。