提出更通用的随机平滑方法,无需光滑密度或全支撑即可高效估计非可微函数梯度。
Generalizing Stochastic Smoothing for Differentiation and Gradient Estimation
- 从基础原理出发,放宽对平滑密度和支撑集的要求
- 在排序、最短路径等任务上实现更低方差的梯度估计
- 适用于黑箱非可微函数,适合机器学习中的可微模拟场景
针对非可微函数(如算法、算子、模拟器)的梯度估计问题,传统随机平滑通过在输入上添加具有全支撑的可微分布来实现平滑并导出梯度。本文从基本原理出发,推导出更宽松假设下的随机平滑方法,无需要求可微密度或全支撑,构建了针对非可微黑箱函数 $f:\mathbb{R}^n\to\mathbb{R}^m$ 的统一松弛与梯度估计框架。从三个不同角度设计方差缩减策略。实验对比了6种分布与最多24种方差缩减方法,在可微排序与排名、图上的可微最短路径、姿态估计的可微渲染及可微冷冻电镜模拟(cryo-ET)中验证有效性。
原文摘要 · Abstract (English)
We deal with the problem of gradient estimation for stochastic differentiable relaxations of algorithms, operators, simulators, and other non-differentiable functions. Stochastic smoothing conventionally perturbs the input of a non-differentiable function with a differentiable density distribution with full support, smoothing it and enabling gradient estimation. Our theory starts at first principles to derive stochastic smoothing with reduced assumptions, without requiring a differentiable density nor full support, and we present a general framework for relaxation and gradient estimation of non-differentiable black-box functions $f:\mathbb{R}^n\to\mathbb{R}^m$. We develop variance reduction for gradient estimation from 3 orthogonal perspectives. Empirically, we benchmark 6 distributions and up to 24 variance reduction strategies for differentiable sorting and ranking, differentiable shortest-paths on graphs, differentiable rendering for pose estimation, as well as differentiable cryo-ET simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。