用因果超图解析训练中批大小如何影响泛化效果。
Actionable Interpretability via Causal Hypergraphs: Unravelling Batch Size Effects in Deep Learning
- 构建超图因果模型,捕捉训练动态中的高阶交互关系。
- 小批量通过增加随机性和平坦极小值提升泛化性能。
- 适合关注优化策略可解释性的深度学习研究者。
尽管批大小对视觉任务泛化的影响已有广泛研究,但在图和文本领域其因果机制仍不明确。本文提出基于超图的因果框架HGCNet,利用深度结构因果模型(DSCMs)揭示批大小如何通过梯度噪声、极小值尖锐度和模型复杂度影响泛化。与依赖静态成对依赖关系的现有方法不同,HGCNet采用超图建模训练动态中的高阶交互。结合do-演算,量化批大小干预的直接与中介效应,提供可解释的因果洞察。在引文网络、生物医学文本和电商评论数据集上的实验表明,HGCNet优于GCN、GAT、PI-GNN、BERT和RoBERTa等强基线。分析显示,较小批大小通过增强随机性与形成更平坦极小值,因果性地提升泛化能力,为深度学习训练策略提供可操作的解释性指导。本工作将可解释性置于原则性架构与优化选择的核心位置,超越事后分析。
原文摘要 · Abstract (English)
While the impact of batch size on generalisation is well studied in vision tasks, its causal mechanisms remain underexplored in graph and text domains. We introduce a hypergraph-based causal framework, HGCNet, that leverages deep structural causal models (DSCMs) to uncover how batch size influences generalisation via gradient noise, minima sharpness, and model complexity. Unlike prior approaches based on static pairwise dependencies, HGCNet employs hypergraphs to capture higher-order interactions across training dynamics. Using do-calculus, we quantify direct and mediated effects of batch size interventions, providing interpretable, causally grounded insights into optimisation. Experiments on citation networks, biomedical text, and e-commerce reviews show that HGCNet outperforms strong baselines including GCN, GAT, PI-GNN, BERT, and RoBERTa. Our analysis reveals that smaller batch sizes causally enhance generalisation through increased stochasticity and flatter minima, offering actionable interpretability to guide training strategies in deep learning. This work positions interpretability as a driver of principled architectural and optimisation choices beyond post hoc analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。