为神经微分方程设计连续时间丢弃法,提升泛化与不确定性估计能力。
Continuum Dropout for Neural Differential Equations
- 将丢弃机制建模为连续时间的随机交替过程,实现对NDE的正则化。
- 在时序和图像分类任务中优于现有方法,且预测置信度更准确。
- 适用于需要可信不确定性估计的场景,如医疗、金融等高风险领域。
神经微分方程(NDEs)擅长建模连续时间动态过程,能有效处理不规则观测、缺失值和噪声等问题。尽管具备诸多优势,NDEs 在采用经典丢弃(dropout)正则化方面面临根本性挑战,导致易过拟合。为此,本文提出连续体丢弃(Continuum Dropout),一种基于交替更新过程理论的通用正则化方法。该方法将丢弃的开/关机制建模为连续时间内的活跃(演化)与静止(暂停)状态交替的随机过程,为防止过拟合和增强NDE的泛化能力提供理论依据。此外,连续体丢弃通过测试时蒙特卡洛采样提供结构化的预测不确定性量化框架。大量实验表明,该方法在多种时序和图像分类任务中优于现有正则化策略,同时生成更校准、更可信的概率估计,验证了其在不确定性感知建模中的有效性。
原文摘要 · Abstract (English)
Neural Differential Equations (NDEs) excel at modeling continuous-time dynamics, effectively handling challenges such as irregular observations, missing values, and noise. Despite their advantages, NDEs face a fundamental challenge in adopting dropout, a cornerstone of deep learning regularization, making them susceptible to overfitting. To address this research gap, we introduce Continuum Dropout, a universally applicable regularization technique for NDEs built upon the theory of alternating renewal processes. Continuum Dropout formulates the on-off mechanism of dropout as a stochastic process that alternates between active (evolution) and inactive (paused) states in continuous time. This provides a principled approach to prevent overfitting and enhance the generalization capabilities of NDEs. Moreover, Continuum Dropout offers a structured framework to quantify predictive uncertainty via Monte Carlo sampling at test time. Through extensive experiments, we demonstrate that Continuum Dropout outperforms existing regularization methods for NDEs, achieving superior performance on various time series and image classification tasks. It also yields better-calibrated and more trustworthy probability estimates, highlighting its effectiveness for uncertainty-aware modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。