对比手工知识与学习表征融合效果,揭示何时有用、何时无效。
When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference

- 固定手工知识注入训练,仅2%开销,对比多种学习方法。
- 不同知识可叠加,相同知识则相互替代,强融合反而干扰性能。
- 冻结特征诊断能事后区分结果,但无法提前预测融合成败。
在数据稀缺场景下,融合先验知识与数据驱动学习具有吸引力,但缺乏可控的评估来判断其何时有益、冗余或有害。本文以一个固定的手工知识源(训练时注入的Gabor目标,开销约2%)为基准,对比多种数据驱动方法(SimCLR、SimSiam、DINO、ImageNet迁移、增强、学习型教师),在统一冻结设置下测试13个数据集、9种主干网络、图像数量从150万到1.28百万、尺寸2.5万至224像素、参数量2.5兆至86兆的配置,涵盖超过$ umRuns$次运行,以及分割与检测移植任务。实验显示三种模式反复出现:不同来源的知识可叠加(如先验知识与DeiT增强在注意力主干上组合,使ViT-B/16在224像素下提升+26点,预算翻倍时仍+6.7);同源知识则相互替代,无法超越单一更优方法;在已有充分信息的初始化中强融合会引发干扰,降幅达-15至-17点,弱化辅助权重可消除此影响。单独分析各源的冻结特征可事后区分三类结果,但不能提前预测。基于冻结特征的增益,在实际标签预算下对端到端增益的预测误差小于0.17点,覆盖30个单元和7个数据集;其分解公式Δ = G + readout(base) 在$ umRate ext{%}$测试单元中符号一致,且可提前预判未见主干族的特征增益。
原文摘要 · Abstract (English)
Fusing prior knowledge with data-driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or harms. We benchmark one fixed hand-crafted knowledge source, a pinned bank of Gabor targets injected only during training at $\sim$2\% overhead, against data-driven alternatives (SimCLR, SimSiam, DINO, ImageNet transfer, augmentation, learned teachers) under one frozen recipe with fixed subsets: 13 datasets, 9 backbones, 150 to 1.28M images, 32--224\,px, 2.5M--86M parameters ($\computeCells$ classification configurations over $\computeRuns$ runs, plus segmentation and detection transplants). Across the training-time combinations we measure, three outcomes recur (decision-level fusion differs). Different-\emph{currency} sources can stack: the prior composes with DeiT augmentation on attention backbones and is worth $+26$ points to ViT-B/16 at $224$\,px, $+6.7$ at twice that budget. Same-currency sources substitute: against effective self-supervised pretraining, the combination never usefully exceeds the better single source. Fusing at full strength into an already-informed initialization interferes in proportion to what it carries: ImageNet transfer, $-15$ to $-17$ points, removed by a weaker auxiliary weight. Frozen-feature diagnostics measured on each source alone separate these outcomes retrospectively but do not predict them: a rule built on them calls one of nine unseen pairs. At a practitioner's own label budget, the frozen-feature gain predicts the end-to-end gain to within $0.17$ points across 30 cells and seven datasets; the underlying decomposition, $Δ= G + \readout(\mathrm{base})$, holds in sign on $\auditRate\%$ of testable cells and is called an unseen backbone family's feature gain in advance. The project page is https://amughrabi.github.io/MomentAux.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。