arXiv:2601.20037cs.LGcs.AI2026-01

让表格数据自动发现特征间关系,既准确又可读。

Structural Compositional Function Networks: Interpretable Functional Compositions for Tabular Discovery

  • 用可微分门控机制建模特征间的数学组合关系
  • 在18个基准上显著提升科学与临床数据表现(p<0.05)
  • 结果以人类可读公式呈现,参数量仅为传统模型的1/10到1/20

尽管表格数据在高风险领域广泛应用,传统深度学习架构常难以超越梯度提升决策树的表现,且缺乏科学可解释性。标准神经网络通常将特征视为独立实体,忽略了定义表格分布的内在流形结构依赖。本文提出结构化组合函数网络(StructuralCFN),通过可微分结构先验引入关系感知归纳偏置。StructuralCFN 显式将每个特征建模为其他特征的数学组合,利用可微自适应门控自动发现最优激活机制(如注意力过滤或抑制极性)。该框架支持结构化知识注入,允许领域先验直接融入架构以引导发现。我们在18个基准上进行严格的10折交叉验证,证明在科学与临床数据集(如Blood Transfusion、Ozone、WDBC)上实现统计显著提升(p < 0.05)。此外,StructuralCFN 提供内在符号可解释性:能恢复数据流形的“规律”为人类可读数学表达式,同时保持紧凑参数量(300–2,500),比标准深度基线小10倍至20倍。

原文摘要 · Abstract (English)

Despite the ubiquity of tabular data in high-stakes domains, traditional deep learning architectures often struggle to match the performance of gradient-boosted decision trees while maintaining scientific interpretability. Standard neural networks typically treat features as independent entities, failing to exploit the inherent manifold structural dependencies that define tabular distributions. We propose Structural Compositional Function Networks (StructuralCFN), a novel architecture that imposes a Relation-Aware Inductive Bias via a differentiable structural prior. StructuralCFN explicitly models each feature as a mathematical composition of its counterparts through Differentiable Adaptive Gating, which automatically discovers the optimal activation physics (e.g., attention-style filtering vs. inhibitory polarity) for each relationship. Our framework enables Structured Knowledge Integration, allowing domain-specific relational priors to be injected directly into the architecture to guide discovery. We evaluate StructuralCFN across a rigorous 10-fold cross-validation suite on 18 benchmarks, demonstrating statistically significant improvements (p < 0.05) on scientific and clinical datasets (e.g., Blood Transfusion, Ozone, WDBC). Furthermore, StructuralCFN provides Intrinsic Symbolic Interpretability: it recovers the governing "laws" of the data manifold as human-readable mathematical expressions while maintaining a compact parameter footprint (300--2,500 parameters) that is over an order of magnitude (10x--20x) smaller than standard deep baselines.

表格数据可解释性组合函数结构先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。